Thanks to our customers, Quadric has outgrown its headquarters. Effective June 18, our new address is 270 East Lane, Burlingame, CA 94010.
Chimera GPNPU
Unified NPU IP: CPU Programmability + Accelerator Efficiency
Build on one IP platform. 1 to 6,400 TOPS. Same architecture. Same toolchain. ASIL-ready.
Burn less power. Data stays local. No round trips to separate accelerators.
Ship new models faster. No code partitioning. ONNX compiles. New operators in C++/Python.
Skip the performance surprises. What you simulate is what you get. Cycle-accurate from day one.
Request Datasheet Schedule a Meeting
The Challenge
From Complexity to Simplicity
The Problem
The Multi-IP Mess
Separate NPU, DSP, CPU. Multiple vendors. Months porting models.
Traditional NPU IP Architecture
- AXI PROGRAMMABLE BUT LOW CAPACITY CPU
- CPU DSP
- Memory Network on Chip HIGH CAPACITY BUT HARDWIRED
Multiple IP Blocks
Separate NPU, DSP, and CPU components that must be integrated, debugged, and maintained independently.
Split Toolchains & Codestreams
Each processor requires its own compiler and debugger. AI workloads partitioned across cores.
Fixed-Function NPUs
Hardware optimized for last year's models. New operators require silicon updates.
The Quadric Approach
One Unified Core
Single Core. 100% C++ Programmable. Single Binary.
GPNPU Core Architecture
- AXI System Fabric PROGRAMMABLE AND HIGH CAPACITY
- Instruction Cache Scalar Unit 2D SIMD Processing Elements ALUMACs Mem External DMA L2 Memory + Internal DMA
Single Unified Core
Matrix, vector, and scalar operations in one execution pipeline. No partitioning.
100% C++ Programmable
New operators added via C++ kernels after deployment. Never blocked by silicon.
Single Unified Binary
One codestream, one toolchain, one debug environment. ONNX and C++ merge seamlessly.
Explore the SDK View Benchmarks
Why Chimera
Benefits of Chimera GPNPU
Simplify your SoC design and speed up porting of new AI models.
System Simplicity
Quadric's solution enables hardware developers to instantiate a single core that can handle an entire AI/ML workload plus the typical digital signal processor functions and signal conditioning workloads often intermixed with inference functions. Dealing with a single core drastically simplifies hardware integration and eases performance optimization. System design tasks such as profiling memory usage and estimating system power consumption are greatly simplified.
Programming Simplicity
Quadric's Chimera GPNPU architecture dramatically simplifies software development since matrix, vector, and control code can all be handled in a single code stream. Graph code from common training toolsets (TensorFlow, PyTorch, ONNX formats) is compiled by the Quadric SDK and can be merged with signal processing code written in C++, all compiled into a single code stream running on a single processor core. The entire subsystem can be debugged in a single debug console.
Future-Proof Flexibility
A Chimera GPNPU can run any AI/ML graph that can be captured in ONNX, and anything written in C++. This is incredibly powerful since SoC developers can quickly write code to implement new neural network operators and libraries long after the SoC has been taped out. This eliminates fear of the unknown and dramatically increases a chip's useful life. As ML models continue to evolve, the payoff from this unified architecture helps future-proof chip design cycles.
How It Works
Unified by Design
Accelerator-level performance with full processor flexibility designed from the ground up to address the constantly evolving AI inference deployment challenges facing system on chip (SoC) developers, the Chimera GPNPU family has a simple yet powerful architecture with demonstrated improved matrix-computation performance over the traditional approach.
- Unified Pipeline: Matrix, vector and scalar code in one execution pipeline. No partitioning required.
- Code-Driven: Continuously optimize performance throughout a device's lifecycle via software updates.
- Future-Ready: Runs classic backbones, Transformers, LLMs, and networks not yet invented.
Chimera GPNPU Block Diagram
A hybrid Von Neumann + 2D SIMD architecture that unifies matrix, vector, and scalar operations in a single execution pipeline
The Chimera GPNPU is entirely driven by code, empowering developers to continuously optimize the performance of their models and algorithms throughout the device's lifecycle. That's why it's ideal to run classic backbone networks, today's newest Transformers and Large Language Models, and whatever new networks are invented tomorrow.
Technical Specifications
Key Architectural Features
A hybrid Von Neumann + 2D SIMD architecture optimized for AI/ML inference
Architecture
- Hybrid Von Neumann + 2D SIMD matrix architecture
- 7-stage, in-order pipeline
- 64b instruction word, single issue per clock
- Scalar/vector/matrix instructions intermixed with granular predication
Memory System
- Distributed local register memories (LRM) with data broadcast networks
- L2 data memory (1MB to 16MB) to minimize off-chip DDR access
- Configurable instruction cache (64/128/256K options)
- Configurable AXI interfaces for independent data and instruction access
Processing
- Optimized for INT8 ML inference with optional FP16 and A8W4 MAC
- 32b ALU DSP operations for full C++ compiler support
- Deterministic, non-speculative execution for predictable performance
Power Efficiency
Compiler-driven, fine-grained clock gating, hierarchical memory minimizes power-hungry DDR access, overlapped compute and data movement within matrix array.
Power Efficiency
Memory Optimization = Power Minimization
ML inference is a data & memory movement optimization problem, not a compute efficiency problem.
Multi-Layer Memory Optimization
Chimera Graph Compiler (CGC) manages data movement across the memory hierarchy
Energy Cost
- 32b write from ALU/MAC Register File 1X
- Local Register Memories 2-3X
- L2 Memory 70X
- Off-chip DDR 225X
Key Insight: Compiler optimizations that keep data resident in the Register File or LRM yield significant power savings.
Technology Comparison
GPNPU vs NPU
Understanding the key differences between traditional NPUs and General Purpose NPUs
| Feature | NPU | GPNPU |
|---|---|---|
| Programmability | Fixed-function hardware blocks | Fully programmable processor |
| New Operators | Requires silicon revision | Write as C++ kernels post-deployment |
| Architecture | Requires CPU/DSP companion | Standalone unified core |
| Code Streams | Split across heterogeneous cores | Single unified codestream |
| Matrix Performance | High | High |
Product Portfolio
The Chimera GPNPU Family
From single-core QC-Nano to multi-instance QC-Multi clusters. Fully synthesizable for any process node.
| TOPS RANGE | CORES | MACs | ALUs |
|---|---|---|---|
| 12nm @ 1.0GHz → 3nm @ 1.7GHz | QC-Multi Multi-core / multi-instance | 32 – 6,400 TOPS | 2 – 64K |
| QC-Ultra High-performance single core | 16 – 108 TOPS | 18K – 32K | 1024 |
| QC-Perform Balanced performance | 4 – 28 TOPS | 12K – 8K | 256 |
| QC-Nano Ultra-efficient edge | 1 – 7 TOPS | 15 | 12 – 2K |
TOPS range varies by process node (12nm @ 1.0GHz to 3nm @ 1.7GHz) and configuration.
Ready to Evaluate?
Get the datasheet. Talk to our architects. See the benchmarks.