Thanks to our customers, Quadric has outgrown its headquarters. Effective June 18, our new address is 270 East Lane, Burlingame, CA 94010.
YOURMODELSONE TOOLCHAINUNIFIED CODEEFFICIENT INFERENCECPUCPUDSPDSPCHIMERA POWERED AI CHIP
$ qcc compile model.onnx
$ qrun --target chimera
AI Chips, Simplified.
AI Processor IP + software stack built to keep pace with your model roadmap.
The Problem
Edge AI Silicon Is a Mess.
Multiple accelerator IPs. Multiple toolchains. And NPUs that can't keep up with the models you need to run.
✗ Chip delays. Multi-accelerator IPs, complex integration and validation push tapeout to the right.
✗ Brittle silicon. Fixed-function NPUs target today's operators, and new models may not map.
✗ Fragmented toolchains. One stack per IP block, and you wait on vendors to port your model.
The Solution
Unified Inference. CPU flexibility. Accelerator efficiency.
Licensable Processor IP for end-to-end inference. One toolchain across your chip line.
✓ Future flexibility. Programmable in Python. New operators and model architectures can map.
✓ Silicon efficiency. Low power. Optimized for single-batch latency. Immediate inference for real-time responsiveness.
✓ Software simplicity. One toolchain. One code stream. No partitioning across accelerators.
The Opportunity
Datacenters Have a Ceiling. Devices Don't.
The next billion AI users won't connect to a datacenter. They'll run inference locally—on silicon built for it.
AI in Your Hands
Multimodal AI on PCs and mobile. Voice, gesture, and personalization. Privacy and latency—solved.
Vision at Scale
Edge servers, smart printers, inspection, security, and analytics—at the edge where data lives.
See, Think, Act
Multimodal perception and decision-making. Sensor fusion, navigation, and manipulation.
Safety at the Edge
Vision, radar, and LiDAR processing with ASIL B/D support. From in-cabin monitoring to full ADAS.
Here's how it's built.
For Architects
Unified Inference with GPNPU
Single-batch performance. Lower power. Deterministic timing. Simpler code.
Single-Batch Latency
Optimized for single-batch latency. Immediate inference for real-time responsiveness.
Power Efficiency
No round trips to separate vector units or DSPs. Data stays local, memory bandwidth drops.
Deterministic Timing
Data movement is instruction-encoded—no routing decisions, no contention. Every cycle is predictable.
Simplified Software
Single instruction stream. No partitioning code between NPU, DSP, and CPU. One codebase.
GPNPU Architecture
Chimera SDK
Flexible Development with Chimera SDK
Compile most models. Extend the rest in C++.
Fast Porting
Graph Compiler auto-compiles hundreds of models. Import ONNX from PyTorch or TensorFlow and run.
Self-Service Extensibility
No waiting on Quadric for new ops. Python via ChiPy™ or C++ extensible with LLVM compiler included.
Pre-Silicon Accuracy
Instruction Set Simulator (ISS) gives cycle-accurate performance prediction before tape-out.
Simpler Software
One processor handles pre-processing, inference, and post-processing. No partitioning across NPU, DSP, CPU.