My RTL works. But where are the actual gates?
“The RTL works in simulation. But simulation is only showing us behavior. Where is the actual hardware structure?”
In Section 8, you verified that your SystemVerilog code behaves with 100% mathematical precision. You drove clock pulses, inspected waveforms, and passed automated assertions.
Now you face the fundamental reality: You have written an abstract description of hardware, but you have not yet built the physical hardware itself. How does human-readable RTL turn into millions of physical transistors on silicon?
If the RTL works, haven't we already built the circuit?
When a software developer writes C or Python code, compiling it creates machine instructions executed by an existing CPU. But in chip design, there is no CPU waiting to execute your RTL.
Simulation answers: “Does this description behave correctly?” But simulation does not answer: “What physical gates, standard cells, and wires should exist on silicon?”
What does RTL actually describe?
Think of RTL as an architectural blueprint for a house (Slide 2). The blueprint tells you where walls, doors, and rooms go, and how people move through the space. But it does not supply the actual physical bricks, timber, nails, or glass.
RTL defines Logical Intent (Boolean math), Register State (Flip-flops), and Dataflow. But it completely lacks physical realities like standard cell models, wire lengths, propagation delays, and thermal profiles.
Simulation verifies behavior (“What should it do?”). Synthesis generates structure (“What gates physically exist?”).
The engineer describes mathematical relationships and register transfers:
module alu_core (
input logic a, b, c,
output logic y
);
// Combinational relationship
assign y = (a & b) | c;
endmoduleSimulation executed the code, but no physical transistors or standard cells have been chosen yet!
Section 8 testbenches verified that when a=1, b=1 or c=1, y=1:
- Logical Intent: Boolean math, additions, multiplexing.
- Register State: Edge-triggered flip-flop storage.
- Dataflow: How data streams from register to register.
- Standard Cells: Which specific foundry gates to use.
- Timing Delays: Nanosecond propagation through silicon.
- Spatial Placement & Routing: Wire tracks and die layout.
Is every RTL statement a gate?
A common beginner misconception is thinking synthesis performs a literal one-to-one translation where line 1 becomes gate 1 and line 2 becomes gate 2.
In reality, the synthesis compiler parses the entire behavioral description, identifies intent, and infers appropriate hardware structures:
assign y = a & b;→ Infers physical AND Logic.assign y = sel ? b : a;orif-else→ Infers physical Multiplexers (MUX).always_ff @(posedge clk)→ Infers edge-triggered D Flip-Flops (DFF).assign y = a + b;→ Infers multi-bit Full Adder chains.
Misconception Breaker: Is every RTL statement a gate? Predict what hardware each SystemVerilog construct infers.
assign y = a & b;A continuous boolean assignment combining two single-bit wires.
What if the RTL describes the same behavior in different ways?
Consider these two boolean equations:
(a & b) | (a & c) // 3 gatesa & (b | c) // 2 gates (Factored)Both produce 100% identical truth tables. But Description B saves 33% of silicon gate area! This reveals the true power of synthesis: Synthesis is not just translation; it is mathematical transformation and optimization.
What does “synthesis” actually do?
Logic synthesis executes a rigorous three-step compilation process (Slide 7):
- Parsing & Elaboration: Translates raw Verilog text into an internal Abstract Syntax Tree (AST) and technology-independent generic boolean network.
- Logic Optimization: Applies Boolean algorithms (like ABC) to simplify equations, remove dead logic, and minimize critical path depth.
- Technology Mapping: Binds generic logic nodes to actual standard cells from the foundry library to satisfy timing and area constraints.
Synthesis is transformation, not translation. Watch the optimization compiler eliminate redundant gates and reduce critical path delay.
assign y = a & (b | c);By applying Boolean distributive laws, the synthesis engine factors out the common term `a`, eliminating an entire 2-input AND gate without changing logical behavior.
Where do the actual gates come from?
Where does the tool get real physical gates? Semiconductor foundries (like TSMC, Intel, or SkyWater) provide a Standard-Cell Library (Slide 6).
Like pre-fabricated Lego bricks, each standard cell (NAND2X1, INVX1, MUX2X1, DFFR_X1) is meticulously designed at the transistor level with a fixed physical height and characterized for delay, leakage power, and silicon area across varying voltages and temperatures.
Standard cells are the pre-characterized “Lego bricks” of silicon foundries. Inspect how technology mapping binds logic equations to physical cells.
Every standard cell in a library has the exact same fixed physical height ($2.72\ \mu m$). This allows automated placement tools to tile millions of cells seamlessly in neat horizontal rows across the silicon die!
2-input NAND gate. The most silicon-efficient universal logic gate in CMOS technology.
cell (NAND2X1) {
area : 3.92;
pin(A) { direction: input; }
pin(B) { direction: input; }
pin(Y) { direction: output;
function: "!(A & B)"; }
}Foundries provide .lib (Liberty) files containing exact delay lookup tables for every input transition time and output capacitive load.
What exactly is a netlist?
Once technology mapping is complete, how does the synthesis engine represent the resulting physical circuit? It outputs a Gate-Level Netlist (Slide 8).
A netlist is a purely structural description containing three building blocks:
The library types used (e.g. NAND2X1, FA_X1).
Unique occurrences of cells (e.g. U101, U102).
Physical interconnect wires linking pins together.
A gate-level netlist is a comprehensive structural parts list: which library cells exist, and how their pins are wired together.
FA_X1 U102 (.A(acc_out[0]), .B(op_b[0]), .CI(sel_sub), .S(next_acc[0]), .CO(carry_net0));
DFFR_X1 U103 (.D(next_acc[0]), .CLK(clk), .RST_N(rst_n), .Q(acc_out[0]));// alu_acc_top Gate-Level Netlistmodule alu_acc_top (clk, rst_n, sel_sub, in_data, acc_out);wire carry_net0, carry_net1, op_b0, op_b1;// 1. Cells, 2. Instances, 3. Nets:MUX2X1U101 (.A(in_data[0]),.B(n_inv0),.S(sel_sub),.Y(op_b[0]));FA_X1U102 (.A(acc_out[0]),.B(op_b[0]),.CI(sel_sub),.S(next_acc[0]).CO(carry_net0));DFFR_X1U103 (.D(next_acc[0]),.CLK(clk),.RST_N(rst_n),.Q(acc_out[0]));FA_X1U104 (.A(acc_out[1]),.B(op_b[1]),.CI(carry_net0),.S(next_acc[1]).CO(carry_net1));DFFR_X1U105 (.D(next_acc[1]),.CLK(clk),.RST_N(rst_n),.Q(acc_out[1]));
1-Bit Full Adder computing sum for bit 0 and propagating carry to bit 1.
| Cell Pin | Direction | Connected Net Wire |
|---|---|---|
| A | INPUT | acc_out[0] |
| B | INPUT | op_b[0] |
| CI | INPUT | sel_sub |
| S | OUTPUT | next_acc[0] |
| CO | OUTPUT | carry_net0 |
Can we actually run synthesis?
In modern open-source silicon engineering, the leading tool for logic synthesis is Yosys (Slide 11).
Below, execute real Yosys synthesis passes on our 4-bit accumulator datapath module. Observe how Yosys reports cell counts, combinational vs sequential instances, and total estimated silicon area!
Execute open-source Yosys to compile your 4-bit accumulator into standard-cell gates and analyze the resulting physical metrics.
| Cell Name | Instance Count | Total Area (µm²) | Area % |
|---|---|---|---|
| NAND2X1 | 4 cells | 15.68 µm² | 11.0% |
| NOR2X1 | 2 cells | 7.84 µm² | 5.5% |
| MUX2X1 | 4 cells | 23.52 µm² | 16.5% |
| FA_X1 | 4 cells | 57.44 µm² | 40.2% |
| DFFR_X1 | 4 cells | 52.24 µm² | 36.6% |
“Our circuit works in simulation and synthesized into gates. But how good is the hardware we actually created?”
We now possess a physical gate-level netlist! But as chip architects, we are judged on three critical competing metrics: Power, Performance, and Area (PPA). Is our critical path fast enough for 1 GHz? Does our circuit burn too much battery? In Section 10, we explore the trade-offs of physical hardware quality!