Thanks to our customers, Quadric has outgrown its headquarters. Effective June 18, our new address is 270 East Lane, Burlingame, CA 94010.

Vision Transformers. At the Edge. No Compromises.

Chimera GPNPU delivers full Vision Transformer support—including attention mechanisms, LayerNorm, and GELU—without the operator limitations of legacy NPUs.

Any Attention Any Norm Any Activation No CPU Fallback

Run ViT in DevStudio View Supported Models

ViT-Base Pipeline

Capability Legacy NPUs (2018-2020) Chimera GPNPU
Optimized for Conv2D and pooling Any ML operator
Dataflow Fixed for ResNet-style networks Flexible for any topology
Attention support Limited or none Native multi-head attention
Activations Hardcoded functions Full GELU, LayerNorm support
Unsupported ops CPU fallback required Complete graph execution in Chimera

Why Legacy NPUs Can't Handle Vision Transformers

Vision Transformers have revolutionized computer vision, outperforming CNNs on classification, detection, and segmentation. But their architecture is fundamentally different from what legacy NPUs were designed to accelerate.

Attention Bottleneck

Multi-head self-attention requires matrix multiplications that don't map efficiently to legacy NPU systolic arrays.

Operator Gaps

LayerNorm, GELU, and Softmax often fall back to CPU, destroying performance and power efficiency.

Dynamic Shapes

Variable sequence lengths in transformers break fixed-tensor NPU compilers designed for static workloads.

Memory Bandwidth

Attention's O(n²) memory access pattern overwhelms legacy NPU memory hierarchies built for sequential access.

How Chimera Runs Vision Transformers

Chimera's General-Purpose NPU architecture was designed from the ground up to handle any ML workload—including the unique demands of transformer models.

Complete Operator Coverage

Every ViT operator runs natively on Chimera—no CPU fallback, no operator gaps. Multi-head attention, LayerNorm, GELU, Softmax, and patch embedding all execute on the GPNPU with full hardware acceleration.

Optimized Attention Compilation

Our Chimera Graph Compiler (CGC) automatically recognizes and optimizes transformer attention patterns, mapping them efficiently to Chimera's compute array and memory hierarchy.

Flexible Memory Architecture

Chimera's software-managed memory handles the variable access patterns of attention without the rigidity of hardware-managed caches. Configure OCM size to match your model's requirements.

Supported Vision Transformer Models

Run industry-standard Vision Transformers on Chimera today, with more models coming soon. All available models are ready for immediate evaluation in DevStudio.

ViT-Base

✓ Available Now
86M Params
224×224 Image
16×16 Patch
Use Cases: Image classification, Feature extraction
Try in DevStudio

ViT-Large

✓ Available Now
307M Params
224×224 Image
16×16 Patch
Use Cases: High-accuracy classification, Transfer learning
Try in DevStudio

ViT-Huge

✓ Available Now
632M Params
224×224 Image
14×14 Patch
Use Cases: Maximum accuracy applications
Try in DevStudio

Swin Transformer

◐ Coming Soon
Hierarchical vision transformer with shifted windows
Use Cases: Object detection, Semantic segmentation

BEVFormer

◐ Coming Soon
Bird's-eye-view transformer for 3D perception
Use Cases: Autonomous driving, Robotics

Custom Models

✦ Contact Us
Bring your own Vision Transformer architecture
Use Cases: Proprietary models, Research networks

Built for Transformer Workloads

Every capability Vision Transformers require, accelerated in hardware.

Feature Description
100% Operators All operators in Chimera
0 CPU fallbacks No CPU fallbacks
1 Multi-Head Self-Attention Native hardware support for scaled dot-product attention with configurable head counts and embedding dimensions
2 LayerNorm Acceleration Hardware-accelerated Layer Normalization without CPU fallback. Handles pre-norm and post-norm architectures
3 GELU & Activation Functions Full support for GELU, Softmax, and all standard activation functions used in transformer architectures
4 Patch Embedding Efficient image-to-patch conversion with configurable patch sizes and embedding dimensions
5 Quantization Support INT8 symmetric and asymmetric quantization with Quantization-Aware Training (QAT) flow for accuracy preservation
6 Multi-Core Scaling Scale ViT inference across multiple Chimera cores for higher throughput with data parallelism

Live Demo Available

View cycle-accurate performance metrics, explore the compiled graph, and evaluate Chimera for your application.

Launch ViT Demo in DevStudio