Thanks to our customers, Quadric has outgrown its headquarters.

Effective June 18, our new address is 270 East Lane, Burlingame, CA 94010.

On-Device LLM

Run Any LLM On-Device with Chimera GPNPU

Chimera GPNPU delivers on-device LLM inference with multi-core clusters and chiplet scaling. Licensed by customers building LLM-capable chips today.

0B
Single-Cluster
0+
Tokens/sec
2 - 8
Cores per Cluster
0B
Via Chiplets
Request More Information Try DevStudio

CHIMERA GPNPU

The Chimera GPNPU architecture enables efficient on-device inference for transformer models up to 30B parameters...

LLM Inference on Chimera

Working FPGA demo running at 50MHz—watch QWEN 0.5B inference on Chimera GPNPU in real-time.
Schedule a Demo
Get a personalized walkthrough of LLM capabilities on Chimera

Customer Validated

Chimera GPNPU has been licensed by customers specifically for LLM workloads. Our implementation has been validated on customer emulation platforms, proving production readiness.

1
Customer License

LLM Use Case Selected

Customer licenses Chimera GPNPU IP for LLM-capable chip design

2
Development

Model Port & Optimization

Customer porting of new LLM using SDK toolkit with attention, pre-fill, and KVcache building blocks

3
Validation

Running on Emulator

Full model validated on customer's emulation platform with correct outputs

Proven RTL

The RTL has been validated on real customer emulation environments. De-risked for your design.

Rapid Model Porting

New LLM models compile from ONNX to optimized C++ in weeks. The CGC compiler handles attention layers, position encoding, and quantization.

Future-Proof Architecture

When the next breakthrough LLM is released, port it in software— no silicon respin required. Your chip stays competitive.

Scale From Single-Core to Multi-Die Chiplets

Chimera GPNPU scales seamlessly from single-core edge deployments to multi-die chiplet configurations, supporting LLMs up to 30B parameters.

Proven Performance Across Model Sizes

Performance validated on emulation with INT4 quantization. Results shown for various LLM sizes on multi-core configurations.

Tokens per Second by Model Size

Model Size Tokens/sec
0.5B LLM1-CORE 42.9 tok/s
0.6B LLM1-CORE 37.7 tok/s
1.7B LLM1-CORE 18.7 tok/s
4B LLM1-CORE 9.7 tok/s
8B LLM4-CORE 20 tok/s

Validated Performance
20+tok/sec
8B parameter LLM on 4-core cluster with INT4 weights, validated on customer emulation platform.
~500ms
Time to first token (512 ctx)
4GB
INT4 weight footprint

Quantization Support

  • W8A8: 8-bit weights, 8-bit activations Supported
  • W4A8: 4-bit weights, 8-bit activations Supported
  • INT4: Full INT4 computation Supported

Memory-Bound Optimization

LLM autoregressive inference is memory-bandwidth bound. Chimera's architecture maximizes bandwidth utilization with optimized weight streaming and KV cache management.

Run Any LLM On-Device, Today and Tomorrow

Chimera's programmable architecture supports any transformer-based LLM. When a new model is released, port it in software—no silicon changes required.

Any Transformer LLM

Architecture Support

Native
Chimera supports standard transformer architectures used by modern LLMs. Port any model from ONNX to optimized C++.

QWEN Family

Alibaba
Validated
QWEN 2.5 and QWEN 3 models validated on customer emulation. 8B model running at 20+ tok/sec.

LLaMA Family

Meta
Supported
Meta's foundational LLM family with multi-core implementation available.

Your Model

Custom
Portable
Bring your own model. CGC compiler handles ONNX conversion to optimized C++ automatically.

Transformer Architecture Support

Core architectural features supported for modern LLMs

Grouped Query Attention

GQA support for efficient key/value sharing across query heads

RoPE Position Encoding

Rotary position embeddings via optimized custom operations

KV Cache Optimization

Efficient autoregressive decoding with optimized cache management

Large Vocabularies

Support for 150K+ token vocabularies with efficient gather operations

Build Your LLM-Capable Chip Today

Whether you're designing edge AI devices, automotive systems, or consumer electronics, Chimera GPNPU delivers the on-device LLM inference your customers demand.

Talk to Our Team

Get detailed documentation, discuss your use case, and learn how Chimera can accelerate your LLM-enabled product roadmap.
Request More Information