Thanks to our customers, Quadric has outgrown its headquarters. Effective June 18, our new address is 270 East Lane, Burlingame, CA 94010.

Chimera GPNPU

Unified NPU IP: CPU Programmability + Accelerator Efficiency

Build on one IP platform. 1 to 6,400 TOPS. Same architecture. Same toolchain. ASIL-ready.

Burn less power. Data stays local. No round trips to separate accelerators.

Ship new models faster. No code partitioning. ONNX compiles. New operators in C++/Python.

Skip the performance surprises. What you simulate is what you get. Cycle-accurate from day one.

Request Datasheet Schedule a Meeting

The Challenge

From Complexity to Simplicity

The Problem

The Multi-IP Mess

Separate NPU, DSP, CPU. Multiple vendors. Months porting models.

Traditional NPU IP Architecture

  • AXI PROGRAMMABLE BUT LOW CAPACITY CPU
  • CPU DSP
  • Memory Network on Chip HIGH CAPACITY BUT HARDWIRED
  1. Multiple IP Blocks

    Separate NPU, DSP, and CPU components that must be integrated, debugged, and maintained independently.

  2. Split Toolchains & Codestreams

    Each processor requires its own compiler and debugger. AI workloads partitioned across cores.

  3. Fixed-Function NPUs

    Hardware optimized for last year's models. New operators require silicon updates.

The Quadric Approach

One Unified Core

Single Core. 100% C++ Programmable. Single Binary.

GPNPU Core Architecture
  • AXI System Fabric PROGRAMMABLE AND HIGH CAPACITY
  • Instruction Cache Scalar Unit 2D SIMD Processing Elements ALUMACs Mem External DMA L2 Memory + Internal DMA
  1. Single Unified Core

    Matrix, vector, and scalar operations in one execution pipeline. No partitioning.

  2. 100% C++ Programmable

    New operators added via C++ kernels after deployment. Never blocked by silicon.

  3. Single Unified Binary

    One codestream, one toolchain, one debug environment. ONNX and C++ merge seamlessly.

Explore the SDK View Benchmarks

Why Chimera

Benefits of Chimera GPNPU

Simplify your SoC design and speed up porting of new AI models.

System Simplicity

Quadric's solution enables hardware developers to instantiate a single core that can handle an entire AI/ML workload plus the typical digital signal processor functions and signal conditioning workloads often intermixed with inference functions. Dealing with a single core drastically simplifies hardware integration and eases performance optimization. System design tasks such as profiling memory usage and estimating system power consumption are greatly simplified.

Programming Simplicity

Quadric's Chimera GPNPU architecture dramatically simplifies software development since matrix, vector, and control code can all be handled in a single code stream. Graph code from common training toolsets (TensorFlow, PyTorch, ONNX formats) is compiled by the Quadric SDK and can be merged with signal processing code written in C++, all compiled into a single code stream running on a single processor core. The entire subsystem can be debugged in a single debug console.

Future-Proof Flexibility

A Chimera GPNPU can run any AI/ML graph that can be captured in ONNX, and anything written in C++. This is incredibly powerful since SoC developers can quickly write code to implement new neural network operators and libraries long after the SoC has been taped out. This eliminates fear of the unknown and dramatically increases a chip's useful life. As ML models continue to evolve, the payoff from this unified architecture helps future-proof chip design cycles.

How It Works

Unified by Design

Accelerator-level performance with full processor flexibility designed from the ground up to address the constantly evolving AI inference deployment challenges facing system on chip (SoC) developers, the Chimera GPNPU family has a simple yet powerful architecture with demonstrated improved matrix-computation performance over the traditional approach.

  • Unified Pipeline: Matrix, vector and scalar code in one execution pipeline. No partitioning required.
  • Code-Driven: Continuously optimize performance throughout a device's lifecycle via software updates.
  • Future-Ready: Runs classic backbones, Transformers, LLMs, and networks not yet invented.

Chimera GPNPU Block Diagram

A hybrid Von Neumann + 2D SIMD architecture that unifies matrix, vector, and scalar operations in a single execution pipeline

The Chimera GPNPU is entirely driven by code, empowering developers to continuously optimize the performance of their models and algorithms throughout the device's lifecycle. That's why it's ideal to run classic backbone networks, today's newest Transformers and Large Language Models, and whatever new networks are invented tomorrow.

Technical Specifications

Key Architectural Features

A hybrid Von Neumann + 2D SIMD architecture optimized for AI/ML inference

Architecture

  • Hybrid Von Neumann + 2D SIMD matrix architecture
  • 7-stage, in-order pipeline
  • 64b instruction word, single issue per clock
  • Scalar/vector/matrix instructions intermixed with granular predication

Memory System

  • Distributed local register memories (LRM) with data broadcast networks
  • L2 data memory (1MB to 16MB) to minimize off-chip DDR access
  • Configurable instruction cache (64/128/256K options)
  • Configurable AXI interfaces for independent data and instruction access

Processing

  • Optimized for INT8 ML inference with optional FP16 and A8W4 MAC
  • 32b ALU DSP operations for full C++ compiler support
  • Deterministic, non-speculative execution for predictable performance

Power Efficiency

Compiler-driven, fine-grained clock gating, hierarchical memory minimizes power-hungry DDR access, overlapped compute and data movement within matrix array.

Power Efficiency

Memory Optimization = Power Minimization

ML inference is a data & memory movement optimization problem, not a compute efficiency problem.

Multi-Layer Memory Optimization

Chimera Graph Compiler (CGC) manages data movement across the memory hierarchy

Energy Cost

  • 32b write from ALU/MAC Register File 1X
  • Local Register Memories 2-3X
  • L2 Memory 70X
  • Off-chip DDR 225X

Key Insight: Compiler optimizations that keep data resident in the Register File or LRM yield significant power savings.

Technology Comparison

GPNPU vs NPU

Understanding the key differences between traditional NPUs and General Purpose NPUs

Feature NPU GPNPU
Programmability Fixed-function hardware blocks Fully programmable processor
New Operators Requires silicon revision Write as C++ kernels post-deployment
Architecture Requires CPU/DSP companion Standalone unified core
Code Streams Split across heterogeneous cores Single unified codestream
Matrix Performance High High

Product Portfolio

The Chimera GPNPU Family

From single-core QC-Nano to multi-instance QC-Multi clusters. Fully synthesizable for any process node.

TOPS RANGE CORES MACs ALUs
12nm @ 1.0GHz → 3nm @ 1.7GHz QC-Multi Multi-core / multi-instance 32 – 6,400 TOPS 2 – 64K
QC-Ultra High-performance single core 16 – 108 TOPS 18K – 32K 1024
QC-Perform Balanced performance 4 – 28 TOPS 12K – 8K 256
QC-Nano Ultra-efficient edge 1 – 7 TOPS 15 12 – 2K

TOPS range varies by process node (12nm @ 1.0GHz to 3nm @ 1.7GHz) and configuration.

Ready to Evaluate?

Get the datasheet. Talk to our architects. See the benchmarks.