ChipFACTORYFoundations
Apply now
Foundations/STAGE 11
Foundations Track/STAGE 11
45 min Physical Design
SECTION 11

The gates need to become a physical chip

The Spatial Transformation

“We have a synthesized netlist and we evaluated its PPA. But gates exist only as a logical abstraction in software. Where do all of these gates physically go on silicon?”

In Section 9 and 10, you synthesized abstract RTL into standard cells and evaluated its Power, Performance, and Area. The circuit is logically complete.

Now comes the great leap into geometry: Physical Design (Place & Route). You must define the physical die, snap standard cells into rows, balance clock distribution trees, and weave kilometers of microscopic copper interconnects across 10 to 15 vertical metal layers!

The Physical Design Pipeline:
FLOORPLAN→PLACEMENT→CTS→ROUTING
11.1•Spatial Reality

A netlist has no address

A synthesized gate netlist describes logical connections (e.g. NAND2 U1 (.A(in), .Y(n1));), but it contains zero spatial information. It does not state where any gate lives in 2D space.

If connected gates are placed far apart on opposite corners of the silicon die, the long connecting wire introduces parasitic resistance ($R$) and capacitance ($C$), turning an instant logical connection into a slow, energy-draining bottleneck!

Interactive Workbench • Slide 1 & 2

A Netlist Has No Address: Logical Connectivity vs. Physical Geometry

A synthesized Verilog netlist states what connects to what, but has no concept of 2D physical coordinates ($X, Y$). Physical design is the art of mapping dimensionless logic nodes into spatial silicon geometry.

Silicon Die Area (500µm × 500µm)Coordinates: (0,0) → (500,500)
DIE BOUNDARYCORE AREALOGICAL NETLIST CLOUD(X, Y Coordinates Undefined)
Physical Impact of Spatial Location:
Total Interconnect Wirelength:---
Interconnect Parasitic Delay:---
Parasitic Capacitance:---
Max Safe Operating Clock:---
The Physical Reality:Gates exist only as logical equations in netlist text. Coordinates are undefined.
11.2•The Spatial Blueprint

What is a floorplan?

Before placing thousands of gates, the physical design engineer must establish the macroscopic blueprint of the chip:

  1. Die Boundary: The full physical dimensions of the silicon chip harvested from the wafer.
  2. Core Area: The internal rectangular region reserved for standard cell logic and macros.
  3. I/O Ring & Power Grid: Peripheral wire-bond pads and VDD/VSS power rings.
  4. Macro Placement: Positioning large pre-designed blocks (such as SRAM memory caches).

Engineers target a Core Utilization of 60%–80%. Over-utilization (>85%) chokes routing tracks; under-utilization (<55%) wastes expensive silicon area.

Interactive Workbench • Slide 3 & 4

Floorplanning: Die vs. Core, Utilization Targets & Macro Placement

Floorplanning establishes the spatial master plan: sizing the Die, bounding the internal Core, placing large SRAM Macros, and balancing Core Utilization (60%–80%) so routing channels never choke.

Core Utilization Target:
70%(Optimal 60%–80%)
50% (Under-utilized)70% (Target Sweet Spot)90% (Over-utilized / Unroutable)
SRAM Memory Macro Placement:
Die Floorplan Schematic (817µm × 817µm)Die: 0.667 mm²
DIE BOUNDARY (817µm)PERIPHERAL I/O PAD RING & POWER RINGS (VDD/VSS)CORE REGION (717µm × 717µm)SRAM MACRO512KB CacheStandard Cells: 70% density
Floorplanning Design Metrics:
Standard Cell + Macro Area:360,000 µm²
Core Area (70% util):514,286 µm²
Routing Whitespace:30% reserved tracks
Total Silicon Die Area:0.667 mm²
Macro Strategy:Corner Aligned (Optimal)
Routable Floorplan Target

60%–80% utilization leaves plenty of routing channels for the router to connect standard cells without shorts.

11.3•Standard Cell Placement

Where should each cell go?

Once the floorplan is locked, millions of standard cells must be assigned physical $(X, Y)$ coordinates. Standard cells have uniform, fixed heights that snap cleanly into predefined horizontal rows like books on a bookshelf.

Placement proceeds in two sequential phases:

  • Phase 1: Global Placement: Solves a continuous hypergraph optimization to group connected gates together and minimize wirelength, ignoring minor cell overlaps.
  • Phase 2: Detailed Legalization: Snaps cells into legal standard cell row boundaries, removes overlaps, and aligns continuous power/ground rails.
Interactive Workbench • Slide 5, 6, 7 & 8

Cell Placement: Global Coarse Clustering vs. Detailed Row Legalization

Standard cell placement operates in two sequential phases: Global Placement clusters connected gates to minimize wirelength while ignoring cell overlaps; Detailed Placement (Legalization) snaps gates into predefined standard cell rows, aligning continuous power/ground rails.

Global Placement (Continuous / Overlapping Cells)COARSE ESTIMATE
VDD RailVSS Rail (Row 1)VDD Rail (Row 2)VSS Rail (Row 3)VDD Rail (Row 4)NAND2OVERLAP!DFF_REGMUX2X1ADDER_B1• Coarse hypergraph positioning (Cell boundaries overlap) •
Placement Engine Mechanics:
01. Global Placement (Slide 7):Solves a mathematical hypergraph partitioning equation to minimize total estimated wirelength across the chip, deliberately ignoring minor geometric cell overlaps.
02. Detailed Legalization (Slide 7):Snaps every cell into strict horizontal rows with uniform cell heights, shifting overlapping cells sideways to guarantee zero geometric design rule violations.
03. Congestion-Driven Placement (Slide 8):If the routing heatmap detects severe track saturation, the placement engine spreads cells apart to create routing whitespace for the metal wires.
11.4•Routing Congestion

What happens to all those wires?

If thousands of standard cells are clustered too tightly, the number of required interconnect wires exceeds the physical capacity of available metal tracks. This triggers Routing Congestion.

Placement engines run background congestion estimation heatmaps. If a region exceeds 85% track capacity, the tool intentionally spreads cells apart, balancing timing-driven placement against congestion-driven placement.

11.5•The Clock Distribution Dilemma

The clock is special

In synchronous digital chips, the clock signal is the master conductor. It must arrive at every register on silicon at precisely the same instant.

Routing a single continuous wire from the clock pad creates massive physical bottlenecks:

  • RC Waveform Degradation: Long resistive wires round sharp clock transitions into slow slopes.
  • Clock Skew: Registers closer to the clock pin receive clock edges earlier than registers far away.
  • Hold-Time Violations: Excessive skew causes race conditions where new data overwrites existing state before it can be safely captured!
Interactive Workbench • Slide 9 & 10

The Clock Distribution Dilemma & Clock Tree Synthesis (CTS)

In synchronous silicon, the clock must arrive at thousands of registers simultaneously. Routing a single continuous wire causes severe RC edge degradation and massive Clock Skew. CTS builds a balanced H-Tree topology with inserted buffers to ensure simultaneous clock arrival.

Recursive 4-Quadrant Symmetrical H-TreeSKEW BALANCED
CLK_ROOTFF1.20nsFF1.22nsFF1.21nsFF1.20ns• Balanced H-Tree: Path Lengths Equal • Clock Skew ≤ 20 ps •
Clock Network Timing Telemetry:
Clock Tree Insertion Delay:1.22 ns
Maximum Clock Skew (Δt):0.02 ns (20 ps)
Clock Transition Slew:0.06 ns (Sharp Buffered Edges)
Hold-Time Violation Risk:SAFE (Zero Hold Violations)
Why Clock Skew is Fatal (Slide 9 & 10):

If the clock reaches Register 2 earlier than Register 1, new data launched from Register 1 can race through short combinational logic and corrupt the contents of Register 2 before it has safely captured its intended value!

11.6•Clock Tree Synthesis (CTS)

How do we build the clock network?

Clock Tree Synthesis (CTS) solves the distribution dilemma by building a symmetrical distribution tree across the die.

The classic model is the recursive H-Tree, which splits the chip area into equal quadrants so the physical distance from the central clock root to every leaf register is identical. Specialized clock buffer inverters are inserted at branching points to regenerate sharp rising and falling edges.

11.7•Multi-Layer Routing & Crosstalk

How do we connect everything?

Once cells are placed and the clock tree is synthesized, interconnect routing weaves electrical connections between all cell pins across 10 to 15 vertical metal layers (M1–M10).

Routing executes in two distinct phases:

  1. Global Routing: Divides the chip into coarse routing tiles and plans macro-level paths to avoid congestion bottlenecks without drawing exact wires.
  2. Detailed Routing: Assigns every wire to an exact physical track on a specific metal layer, inserting vertical tungsten vias to transition between layers while obeying thousands of foundry Design Rule Check (DRC) constraints.
Interactive Workbench • Slide 11, 12, 13 & 14

3D Multi-Layer Metal Routing, Interconnect RC & Crosstalk Noise

Modern semiconductor processes feature vertical stacks of 10 to 15 metal layers (M1–M10). Lower metals route dense, local standard cells; thick upper superhighways deliver power grids and global clock signals. Microscopic vias connect between layers.

16nm Process Metal Stack Cross-Section:METAL 6 (M6) • GLOBAL POWER GRID (VDD/VSS)METAL 5 (M5) • GLOBAL CLOCK DISTRIBUTION H-TREEVia 4METAL 4 (M4) • Inter-Block Data Buses (Medium Pitch)METAL 3 (M3) • Datapath RoutingVia 2METAL 2 (M2) • Dense Standard Cell InterconnectsMETAL 1 (M1) • Local Standard Cell Pin ConnectionsACTIVE SILICON SUBSTRATE (FinFET Transistors)Source / Drain / Gate Poly (FEOL)
The Metal Layer Hierarchy (Slide 12):
Lower Metals (M1–M3):Physically ultra-thin with tight pitch. High electrical resistance. Strictly reserved for short, local connections inside standard cell rows.
Intermediate Metals (M4–M5):Moderate thickness. Utilized for long signal routing across different functional blocks.
Top Superhighways (M6+):Thick, ultra-low resistance copper layers. Reserved for power grids (VDD/VSS) and global clock distribution.
11.8•OpenROAD Automated Flow

What does OpenROAD actually do?

In modern open-source silicon engineering, the leading toolchain for physical implementation is OpenROAD (Slide 15).

Below, execute the full autonomous physical design flow on our 4-bit accumulator design. Watch the netlist initialize a floorplan, legalize standard cells into rows, synthesize an H-Tree clock distribution network, and route all metal interconnections into clean GDSII polygons!

Interactive Workbench • Slide 15 & 16

The OpenROAD Interactive Physical Design Pipeline

OpenROAD is the open-source EDA framework that takes your synthesized gate-level netlist and generates an automated, timing-clean, manufacturing-ready GDSII silicon layout.

01 / FLOORPLANDie & Power Grid70% core utilization
02 / PLACERow Legalization32 standard cells placed
03 / CTSClock Tree SynthesisSkew ≤ 18 ps balanced
04 / ROUTEGlobal & Detailed Route0 DRC wire violations
Design Target: alu_acc_top.sv • Sky130 PDK
Layout Viewer (OpenROAD GUI Representation)GDSII READY
OpenROAD Complete: Clean Layout-Ready GDSII Silicon Output
openroad_flow.logSTATUS: OK
[INFO ODB-0022] Reading LEF file: sky130_fd_sc_hd.tlef
[INFO ODB-0023] Initializing floorplan: 480um x 480um
[INFO IFP-0001] Core utilization set to 70.0%
[INFO GPL-0002] Global placement wirelength: 1,420 um
[INFO DPL-0001] Legalized 32 standard cells (0 overlaps)
[INFO CTS-0001] Synthesizing clock tree: 4 leaf registers
[INFO CTS-0004] Calculated clock skew: 18.2 ps (MET)
[INFO GRL-0001] Global routing completed (0 congestion)
[INFO DRL-0001] Detailed routing: 100% routed on M1-M4
[INFO DRC-0000] Final design rule check: 0 violations
[SUCCESS] GDSII Layout exported successfully.
Physical Implementation Complete

Your SystemVerilog RTL has traversed the entire ASIC pipeline: text → netlist → floorplan → placement → CTS → routing → silicon polygons!

SECTION 11 MASTER RECAP

From Logical Gates to a Physical Silicon Layout

You have taken your digital circuit all the way through physical design: defining floorplan boundaries, legalizing cells into rows, synthesizing a balanced clock tree, and routing multi-layer metal tracks. The design is now a complete geometric layout!

The Next Engineering Frontier • Section 12

How do we know the physical chip is actually going to work?

The layout looks complete and beautiful. But how do we prove that none of the millions of microscopic metal polygons violate foundry manufacturing rules? How do we verify that wire resistance didn't destroy timing margins? In Section 12, we enter Physical Verification: Static Timing Analysis (STA), Design Rule Checking (DRC), and Layout Versus Schematic (LVS)!