Gate-level simulation on a GPU
Gate-level simulation for billion-gate designs, on one GPU.
Xerolane simulates synthesized gate-level netlists on one GPU, running many vectors at the same time. 640,000 vectors through a 16.2-million-gate design take 16.6 seconds on a laptop.
640,000 vectors through 16.2M gates (EPFL sixteen), after 3.1 s of elaboration. 8 GB laptop GPU.
1,000,000 vectors through 93.4M gates: 19.2 s to elaborate, 182.6 s to simulate. Same laptop.
200M gates, 16 vectors × 5,000 cycles, and a SAIF file covering all 200M nets. Same laptop.
tested end to end on one GPU. Two billion run on the 8 GB laptop. The next target is higher.
The simulator
Xerolane evaluates the synthesized netlist, gates and flip-flops, cycle by cycle. It sits after synthesis, where every design is re-verified for correctness, reset and X behaviour, coverage and sign-off, and where runtime is the bottleneck.
A sweep costs close to one run
Independent vectors simulate together on one GPU. How many run at once is a runtime setting.
0, 1, X and Z
Full 4-state for reset, initialization and contention. A faster 2-state path for volume regression. Several clock domains, both edges.
Simulate many
Elaboration is cached against the netlist and build, so later runs of the same design go straight to simulation.
Measured
Laptop figures are from an RTX 4060 with 8 GB, 24 CPU cores and 64 GB RAM, a machine under $2,000. EPFL designs are from the public EPFL Combinational Benchmark Suite; the enlarged ones use ABC's double.
Time for a workload, by design size
| Design | Gates | Workload | Elaboration | Simulation | Machine |
|---|---|---|---|---|---|
| Synthetic, sequential | 2,000,000 | 10,000 cycles | 0.28 s | 3.7 s | Laptop |
| EPFL sixteen | 16,216,836 | 640,000 vectors | 3.1 s | 16.6 s | Laptop |
| EPFL sixteen | 16,216,836 | 640,000 vectors | 5.4 s | 6.8 s | L40, rented pod |
| EPFL twentythree ×2 | 46,680,504 | 1,000,000 vectors | 9.3 s | 88.0 s | Laptop |
| EPFL twentythree ×4 | 93,360,248 | 1,000,000 vectors | 19.2 s | 182.6 s | Laptop |
| Synthetic, 20M flip-flops | 200,000,128 | 16 vectors × 5,000 cycles, SAIF of every net | 42 s | 15 min 56 send to end | Laptop |
Capacity, by card
| Card | Largest design run | GPU memory used |
|---|---|---|
| RTX 4060 laptop, 8 GB | 2 billion gates, 2-state and 4-state | 1B gates in 7.7 GB |
| RTX 4090, 24 GB | 500 million gates | 12.6 GB |
| NVIDIA L40, 48 GB | 4.5 billion gates, end to end | — |
| NVIDIA B200, 183 GB | 200 million gates, many vectors | 15.8 GB |
Correctness: byte-identical to independent reference simulators
| Design or suite | Scope | Result |
|---|---|---|
| A 736K-gate design with RAMs | Compared with an open-source GPU gate-level simulator | Byte-identical |
| Vortex, MIT RISC-V GPGPU | 2.25M gates, 64 vector sets of 10,000 cycles | 64 of 64 byte-identical |
| EPFL sixteen | 16.2M gates, 640,000 vectors | Every vector exact |
| LogikBench RTL corpus | 159 designs, from RTL through Yosys | Byte-identical |
| ISCAS'85/'89, EPFL, LGSynth, MCNC | 2-state and 4-state | No mismatches |
What it does
Every flow runs on the same engine, the same netlist and the same vectors.
- Multiple clock domainsSeveral clocks, per-flop clock assignment, both edges.
- Timed simulationPer-arc rise and fall delays, inertial pulse rejection, SDF back-annotation, VCD output.
- Toggle coverageRecorded during simulation, at about 5% overhead.
- Toggle counts and SAIFEvery net's full transition counts, written as SAIF for power analysis. Counting adds 15% to simulation time.
- Save and restoreThree ways back in: replay, restart with fresh stimulus, or one saved state fanned out across many vectors. Coverage survives the restore.
- Full state dumpEvery flip-flop and net, at any cycle.
- DebugFind the first cycle and output where two runs part, trace an X back to where it came from, and export VCD for any window of cycles.
- Stuck-at fault simulationCombinational and full-scan, many faults at once, with fault collapsing.
- Test pattern generation (ATPG)Generates and compacts test vectors on the same engine: 99.27% test coverage over 6.3 million faults on a 2.25M-gate RISC-V GPU core.
- Design inputVerilog RTL through Yosys, gate-level Verilog, AIG and
.bench.
Running it
Xerolane installs as a set of binaries on a Linux machine with one GPU, NVIDIA or AMD. Both produce identical traces on the same design and stimulus. No compiler or GPU toolkit is needed on your side.
Your netlist stays on your machine.
Try it on your design
Send a netlist, or tell us its size and what you run on it today. We will come back with elaboration and simulation times for your workload.
Company
Xerolane Design Systems LLP is based in Noida, India.
Manu Lauria, founder
An IIT alumnus and former Cadence engineer with 24 years in EDA. He led the NCSim India team and then Xcelium R&D Operations.
Contact
Evaluations, partnerships and design data.