Thanks to our customers, Quadric has outgrown its headquarters. Effective June 18, our new address is 270 East Lane, Burlingame, CA 94010.

YOURMODELSONE TOOLCHAINUNIFIED CODEEFFICIENT INFERENCECPUCPUDSPDSPCHIMERA POWERED AI CHIP

$ qcc compile model.onnx

$ qrun --target chimera

AI Chips, Simplified.

AI Processor IP + software stack built to keep pace with your model roadmap.

The Problem

Edge AI Silicon Is a Mess.

Multiple accelerator IPs. Multiple toolchains. And NPUs that can't keep up with the models you need to run.

✗ Chip delays. Multi-accelerator IPs, complex integration and validation push tapeout to the right.

✗ Brittle silicon. Fixed-function NPUs target today's operators, and new models may not map.

✗ Fragmented toolchains. One stack per IP block, and you wait on vendors to port your model.

The Solution

Unified Inference. CPU flexibility. Accelerator efficiency.

Licensable Processor IP for end-to-end inference. One toolchain across your chip line.

✓ Future flexibility. Programmable in Python. New operators and model architectures can map.

✓ Silicon efficiency. Low power. Optimized for single-batch latency. Immediate inference for real-time responsiveness.

✓ Software simplicity. One toolchain. One code stream. No partitioning across accelerators.

The Opportunity

Datacenters Have a Ceiling. Devices Don't.

The next billion AI users won't connect to a datacenter. They'll run inference locally—on silicon built for it.

AI in Your Hands

Multimodal AI on PCs and mobile. Voice, gesture, and personalization. Privacy and latency—solved.

Vision at Scale

Edge servers, smart printers, inspection, security, and analytics—at the edge where data lives.

See, Think, Act

Multimodal perception and decision-making. Sensor fusion, navigation, and manipulation.

Safety at the Edge

Vision, radar, and LiDAR processing with ASIL B/D support. From in-cabin monitoring to full ADAS.

Here's how it's built.

For Architects

Unified Inference with GPNPU

Single-batch performance. Lower power. Deterministic timing. Simpler code.

Single-Batch Latency

Optimized for single-batch latency. Immediate inference for real-time responsiveness.

Power Efficiency

No round trips to separate vector units or DSPs. Data stays local, memory bandwidth drops.

Deterministic Timing

Data movement is instruction-encoded—no routing decisions, no contention. Every cycle is predictable.

Simplified Software

Single instruction stream. No partitioning code between NPU, DSP, and CPU. One codebase.

GPNPU Architecture

Chimera SDK

Flexible Development with Chimera SDK

Compile most models. Extend the rest in C++.

Fast Porting

Graph Compiler auto-compiles hundreds of models. Import ONNX from PyTorch or TensorFlow and run.

Self-Service Extensibility

No waiting on Quadric for new ops. Python via ChiPy™ or C++ extensible with LLVM compiler included.

Pre-Silicon Accuracy

Instruction Set Simulator (ISS) gives cycle-accurate performance prediction before tape-out.

Simpler Software

One processor handles pre-processing, inference, and post-processing. No partitioning across NPU, DSP, CPU.