Thanks to our customers, Quadric has outgrown its headquarters.
Effective June 18, our new address is 270 East Lane, Burlingame, CA 94010.
On-Device LLM
Run Any LLM On-Device with Chimera GPNPU
Chimera GPNPU delivers on-device LLM inference with multi-core clusters and chiplet scaling. Licensed by customers building LLM-capable chips today.
0B
Single-Cluster
0+
Tokens/sec
2 - 8
Cores per Cluster
0B
Via Chiplets
Request More Information Try DevStudio
CHIMERA GPNPU
The Chimera GPNPU architecture enables efficient on-device inference for transformer models up to 30B parameters...
LLM Inference on Chimera
Working FPGA demo running at 50MHz—watch QWEN 0.5B inference on Chimera GPNPU in real-time.
Schedule a Demo
Get a personalized walkthrough of LLM capabilities on Chimera
Customer Validated
Chimera GPNPU has been licensed by customers specifically for LLM workloads. Our implementation has been validated on customer emulation platforms, proving production readiness.
1
Customer License
LLM Use Case Selected
Customer licenses Chimera GPNPU IP for LLM-capable chip design
2
Development
Model Port & Optimization
Customer porting of new LLM using SDK toolkit with attention, pre-fill, and KVcache building blocks
3
Validation
Running on Emulator
Full model validated on customer's emulation platform with correct outputs
Proven RTL
The RTL has been validated on real customer emulation environments. De-risked for your design.
Rapid Model Porting
New LLM models compile from ONNX to optimized C++ in weeks. The CGC compiler handles attention layers, position encoding, and quantization.
Future-Proof Architecture
When the next breakthrough LLM is released, port it in software— no silicon respin required. Your chip stays competitive.
Scale From Single-Core to Multi-Die Chiplets
Chimera GPNPU scales seamlessly from single-core edge deployments to multi-die chiplet configurations, supporting LLMs up to 30B parameters.
Proven Performance Across Model Sizes
Performance validated on emulation with INT4 quantization. Results shown for various LLM sizes on multi-core configurations.
Tokens per Second by Model Size
| Model Size | Tokens/sec |
|---|---|
| 0.5B LLM1-CORE | 42.9 tok/s |
| 0.6B LLM1-CORE | 37.7 tok/s |
| 1.7B LLM1-CORE | 18.7 tok/s |
| 4B LLM1-CORE | 9.7 tok/s |
| 8B LLM4-CORE | 20 tok/s |
Validated Performance
20+tok/sec
8B parameter LLM on 4-core cluster with INT4 weights, validated on customer emulation platform.
~500ms
Time to first token (512 ctx)
4GB
INT4 weight footprint
Quantization Support
- W8A8: 8-bit weights, 8-bit activations Supported
- W4A8: 4-bit weights, 8-bit activations Supported
- INT4: Full INT4 computation Supported
Memory-Bound Optimization
LLM autoregressive inference is memory-bandwidth bound. Chimera's architecture maximizes bandwidth utilization with optimized weight streaming and KV cache management.
Run Any LLM On-Device, Today and Tomorrow
Chimera's programmable architecture supports any transformer-based LLM. When a new model is released, port it in software—no silicon changes required.
Any Transformer LLM
Architecture Support
Native
Chimera supports standard transformer architectures used by modern LLMs. Port any model from ONNX to optimized C++.
QWEN Family
Alibaba
Validated
QWEN 2.5 and QWEN 3 models validated on customer emulation. 8B model running at 20+ tok/sec.
LLaMA Family
Meta
Supported
Meta's foundational LLM family with multi-core implementation available.
Your Model
Custom
Portable
Bring your own model. CGC compiler handles ONNX conversion to optimized C++ automatically.
Transformer Architecture Support
Core architectural features supported for modern LLMs
Grouped Query Attention
GQA support for efficient key/value sharing across query heads
RoPE Position Encoding
Rotary position embeddings via optimized custom operations
KV Cache Optimization
Efficient autoregressive decoding with optimized cache management
Large Vocabularies
Support for 150K+ token vocabularies with efficient gather operations
Build Your LLM-Capable Chip Today
Whether you're designing edge AI devices, automotive systems, or consumer electronics, Chimera GPNPU delivers the on-device LLM inference your customers demand.
Talk to Our Team
Get detailed documentation, discuss your use case, and learn how Chimera can accelerate your LLM-enabled product roadmap.
Request More Information