The gates need to become a physical chip
“We have a synthesized netlist and we evaluated its PPA. But gates exist only as a logical abstraction in software. Where do all of these gates physically go on silicon?”
In Section 9 and 10, you synthesized abstract RTL into standard cells and evaluated its Power, Performance, and Area. The circuit is logically complete.
Now comes the great leap into geometry: Physical Design (Place & Route). You must define the physical die, snap standard cells into rows, balance clock distribution trees, and weave kilometers of microscopic copper interconnects across 10 to 15 vertical metal layers!
A netlist has no address
A synthesized gate netlist describes logical connections (e.g. NAND2 U1 (.A(in), .Y(n1));), but it contains zero spatial information. It does not state where any gate lives in 2D space.
If connected gates are placed far apart on opposite corners of the silicon die, the long connecting wire introduces parasitic resistance ($R$) and capacitance ($C$), turning an instant logical connection into a slow, energy-draining bottleneck!
A Netlist Has No Address: Logical Connectivity vs. Physical Geometry
A synthesized Verilog netlist states what connects to what, but has no concept of 2D physical coordinates ($X, Y$). Physical design is the art of mapping dimensionless logic nodes into spatial silicon geometry.
What is a floorplan?
Before placing thousands of gates, the physical design engineer must establish the macroscopic blueprint of the chip:
- Die Boundary: The full physical dimensions of the silicon chip harvested from the wafer.
- Core Area: The internal rectangular region reserved for standard cell logic and macros.
- I/O Ring & Power Grid: Peripheral wire-bond pads and VDD/VSS power rings.
- Macro Placement: Positioning large pre-designed blocks (such as SRAM memory caches).
Engineers target a Core Utilization of 60%–80%. Over-utilization (>85%) chokes routing tracks; under-utilization (<55%) wastes expensive silicon area.
Floorplanning: Die vs. Core, Utilization Targets & Macro Placement
Floorplanning establishes the spatial master plan: sizing the Die, bounding the internal Core, placing large SRAM Macros, and balancing Core Utilization (60%–80%) so routing channels never choke.
60%–80% utilization leaves plenty of routing channels for the router to connect standard cells without shorts.
Where should each cell go?
Once the floorplan is locked, millions of standard cells must be assigned physical $(X, Y)$ coordinates. Standard cells have uniform, fixed heights that snap cleanly into predefined horizontal rows like books on a bookshelf.
Placement proceeds in two sequential phases:
- Phase 1: Global Placement: Solves a continuous hypergraph optimization to group connected gates together and minimize wirelength, ignoring minor cell overlaps.
- Phase 2: Detailed Legalization: Snaps cells into legal standard cell row boundaries, removes overlaps, and aligns continuous power/ground rails.
Cell Placement: Global Coarse Clustering vs. Detailed Row Legalization
Standard cell placement operates in two sequential phases: Global Placement clusters connected gates to minimize wirelength while ignoring cell overlaps; Detailed Placement (Legalization) snaps gates into predefined standard cell rows, aligning continuous power/ground rails.
What happens to all those wires?
If thousands of standard cells are clustered too tightly, the number of required interconnect wires exceeds the physical capacity of available metal tracks. This triggers Routing Congestion.
Placement engines run background congestion estimation heatmaps. If a region exceeds 85% track capacity, the tool intentionally spreads cells apart, balancing timing-driven placement against congestion-driven placement.
The clock is special
In synchronous digital chips, the clock signal is the master conductor. It must arrive at every register on silicon at precisely the same instant.
Routing a single continuous wire from the clock pad creates massive physical bottlenecks:
- RC Waveform Degradation: Long resistive wires round sharp clock transitions into slow slopes.
- Clock Skew: Registers closer to the clock pin receive clock edges earlier than registers far away.
- Hold-Time Violations: Excessive skew causes race conditions where new data overwrites existing state before it can be safely captured!
The Clock Distribution Dilemma & Clock Tree Synthesis (CTS)
In synchronous silicon, the clock must arrive at thousands of registers simultaneously. Routing a single continuous wire causes severe RC edge degradation and massive Clock Skew. CTS builds a balanced H-Tree topology with inserted buffers to ensure simultaneous clock arrival.
If the clock reaches Register 2 earlier than Register 1, new data launched from Register 1 can race through short combinational logic and corrupt the contents of Register 2 before it has safely captured its intended value!
How do we build the clock network?
Clock Tree Synthesis (CTS) solves the distribution dilemma by building a symmetrical distribution tree across the die.
The classic model is the recursive H-Tree, which splits the chip area into equal quadrants so the physical distance from the central clock root to every leaf register is identical. Specialized clock buffer inverters are inserted at branching points to regenerate sharp rising and falling edges.
How do we connect everything?
Once cells are placed and the clock tree is synthesized, interconnect routing weaves electrical connections between all cell pins across 10 to 15 vertical metal layers (M1–M10).
Routing executes in two distinct phases:
- Global Routing: Divides the chip into coarse routing tiles and plans macro-level paths to avoid congestion bottlenecks without drawing exact wires.
- Detailed Routing: Assigns every wire to an exact physical track on a specific metal layer, inserting vertical tungsten vias to transition between layers while obeying thousands of foundry Design Rule Check (DRC) constraints.
3D Multi-Layer Metal Routing, Interconnect RC & Crosstalk Noise
Modern semiconductor processes feature vertical stacks of 10 to 15 metal layers (M1–M10). Lower metals route dense, local standard cells; thick upper superhighways deliver power grids and global clock signals. Microscopic vias connect between layers.
What does OpenROAD actually do?
In modern open-source silicon engineering, the leading toolchain for physical implementation is OpenROAD (Slide 15).
Below, execute the full autonomous physical design flow on our 4-bit accumulator design. Watch the netlist initialize a floorplan, legalize standard cells into rows, synthesize an H-Tree clock distribution network, and route all metal interconnections into clean GDSII polygons!
The OpenROAD Interactive Physical Design Pipeline
OpenROAD is the open-source EDA framework that takes your synthesized gate-level netlist and generates an automated, timing-clean, manufacturing-ready GDSII silicon layout.
[INFO ODB-0022] Reading LEF file: sky130_fd_sc_hd.tlef
[INFO ODB-0023] Initializing floorplan: 480um x 480um
[INFO IFP-0001] Core utilization set to 70.0%
[INFO GPL-0002] Global placement wirelength: 1,420 um
[INFO DPL-0001] Legalized 32 standard cells (0 overlaps)
[INFO CTS-0001] Synthesizing clock tree: 4 leaf registers
[INFO CTS-0004] Calculated clock skew: 18.2 ps (MET)
[INFO GRL-0001] Global routing completed (0 congestion)
[INFO DRL-0001] Detailed routing: 100% routed on M1-M4
[INFO DRC-0000] Final design rule check: 0 violations
[SUCCESS] GDSII Layout exported successfully.Your SystemVerilog RTL has traversed the entire ASIC pipeline: text → netlist → floorplan → placement → CTS → routing → silicon polygons!