From research to an algorithm

Relaxation proposes. Exact scoring decides.

A smooth model gives the search a direction. A legal move is kept only when the exact benchmark score improves. After discovery, the placer runs entirely as GPU code.

0.93663mean proxy at 900 s
1.48%below published rank 1
17 / 17legal final placements
48 hresearch campaign
01 / RELAX

Make the grid navigable.

Spread each pin's demand across neighboring cells without changing its total mass. Use gradients to propose global movement.

Approximate guidance
02 / GATE

Keep verified improvements.

Check legality and the exact score. Reject a misleading proposal; retain the incumbent. Incremental GPU scoring makes large batches practical.

Exact acceptance

Comparison with published results. Archgen remains the official challenge winner. Numaro's score is reverified by the campaign's standalone verifier and has not been confirmed by the organizers.

Inspect all 17 circuits 900 s · seed 0 / 300 s · three seeds
Mean proxy cost; lower is better. The 300-second column averages seeds 0, 1, and 2.
CircuitSplit900 s300 s meanLegal
ibm01Development0.755590.76578Yes
ibm02Validation0.913300.92367Yes
ibm03Held out0.859830.86747Yes
ibm04Development0.885700.89643Yes
ibm06Development0.980191.00380Yes
ibm07Development0.941640.96285Yes
ibm08Held out0.983160.99398Yes
ibm09Development0.752330.75752Yes
ibm10Validation0.899210.91127Yes
ibm11Development0.761160.77410Yes
ibm12Development1.066221.07678Yes
ibm13Held out0.803090.82143Yes
ibm14Development1.030931.03845Yes
ibm15Development0.989551.01847Yes
ibm16Validation0.957980.97174Yes
ibm17Development1.163621.16364Yes
ibm18Held out1.179151.16305Yes
MeanAll circuits0.936630.94767All

The frozen method commit is f264d1f8. Held-out circuits are ibm03, ibm08, ibm13, and ibm18.

What happened after routing? Six placements of ariane136

All six placements routed successfully with the same RTL and flow. Lower proxy correlated with lower power; the wirelength trend was not statistically significant, and timing had no relationship with proxy. Total negative slack was zero in all six runs.

Six-placement ariane136 study: total power tracks proxy, routed wirelength has a non-significant trend, and timing does not track proxy.
One design, six placements. Only the power correlation is significant at p < 0.05. This is limited physical validation, not a universal chip-quality claim.

A separate exploratory comparison with OpenROAD RTL-MP improved some routing-demand and cell-density tails but worsened mean congestion and IR drop. The two placements needed different macro halos, so that comparison was not controlled. An earlier timing-improvement claim was retracted after the six-placement study.

Plain language

What this result means

Before a chip's small logic cells are placed and its wires routed, large memory blocks and hard IP need locations on the floorplan. Those choices shape the space available to every later tool. Here the AI performed the research that created a placement algorithm: it proposed hypotheses, wrote code, ran experiments, and recorded what survived. The delivered algorithm uses no language model or trained placement model at run time.

  • The 900-second run lowers the published rank-1 mean from 0.9507 to 0.93663, a 1.48% reduction. The comparison uses reported baseline scores on hardware we do not control.
  • All three 300-second seed means—0.94808, 0.94690, and 0.94803—are below 0.9507. These are complete 17-circuit sweeps, not a selection of the best circuits.
  • Ten development circuits, three validation circuits, and four test circuits were assigned before optimization. The algorithm and parameters were frozen before test evaluation.
  • Twelve search modifications were rejected. The useful changes made the congestion objective smoother for guidance and exact evaluation cheaper for decisions.

Visual notes

How to read the result

The initial and Relax-and-Gate ibm17 floorplans beside horizontal and vertical routing-demand maps. The optimized horizontal peak falls from 3.39 to 1.91.
A floorplan and its routing demand, before and afterOn ibm17, joint optimization of hard macros and cell clusters lowers the proxy from 1.7395 to 1.1636. Both rows use the same demand-to-capacity color scale. These are benchmark demand maps, not measurements from a detailed router.Open full-size figure ↗
Mean benchmark proxy for the three published leaders and Relax-and-Gate, with a separate comparison of the 300- and 900-second search budgets.
The benchmark comparison and two search budgetsOnly complete 17-circuit evaluations appear here. The 300-second result averages three seeds; the 900-second result uses seed 0. Baseline scores are reported values on different hardware. The error bar shows the sample standard deviation of the three complete-run means, not baseline uncertainty.Open full-size figure ↗

Result table

A lower mean proxy cost across all 17 benchmark circuits.

CellBaselineNumaroDeltaNote
Archgen, published rank 10.95070.93663−1.48%900 s/circuit · seed 0 · one L40S
Carrotato / AbuPlace, rank 20.95220.93663−1.64%published baseline mean
JaneRT, rank 30.96940.93663−3.38%published baseline mean
RePlAce1.45780.93663−35.75%same 17-circuit suite
300-second search budget0.95070.94767−0.32%mean across 3 seeds · all 51 placements legal

Method

How it was found

The campaign ran for 48 hours on one L40S. A human chose the problem and compute budget and reviewed the result; the research system generated hypotheses, implemented experiments, and assigned verdicts. The method jointly moves hard macros and cell clusters, first using a differentiable relaxation and then batched discrete search. Both stages keep only exact-scored improvements.

  • Start from the benchmark placement and legalize the hard macros.
  • Replace pin snapping with bilinear mass distribution and route snapping with continuous cell overlap, preserving total demand.
  • Use the relaxed objective to propose global moves, then score legal checkpoints with the exact benchmark objective.
  • Batch discrete proposals on the GPU. Incremental scoring updates the moved macro's affected nets; combined moves are rescored before acceptance.
  • Return the best exact-scored legal placement, then verify its coordinates independently of the search implementation.

Verification

How it was checked

The campaign's standalone verifier imports nothing from the search code. It reloads the original benchmark netlists through the reference TILOS evaluator, recomputes scores from the final coordinates, and checks overlaps, bounds, and fixed macros. The campaign recorded component agreement within 10⁻¹⁴ relative error over 68 full-evaluator comparisons and maximum incremental deviation of 2.2×10⁻¹⁶. For this page, the aggregate and seed means were recalculated from the stored result records; the original GPU campaign and OpenROAD flows were not rerun.

Scope

What is not being claimed

This is a lower reported benchmark proxy, not an official challenge win, an optimality proof, or a general claim of better chips. The four test circuits belong to the same benchmark family. Physical validation is limited to two designs, with a controlled six-placement study on one design whose timing constraint does not bind. Only the power correlation was significant. An exploratory comparison with OpenROAD RTL-MP improves some tail metrics but worsens mean congestion and IR drop, and uses different macro halos. Smooth optimization, exact acceptance, and incremental evaluation also appear in prior work; the paper isolates its differences from AbuPlace. LLM token usage and cost were not instrumented.

References

Baseline sources

Citation

How to cite

Numaro AI Autoresearch. "AI Autoresearch for Chip Design: Autonomous Discovery of a Macro Placement Algorithm." Numaro Research Report NUMARO-2026-022, 2026.

@techreport{numaro2026MacroPlacementRelax,
  title = {AI Autoresearch for Chip Design: Autonomous Discovery of a Macro Placement Algorithm},
  author = {Numaro AI Autoresearch},
  institution = {Numaro},
  number = {NUMARO-2026-022},
  year = {2026},
  url = {https://numaro.tech/research/macro-placement-relax-and-gate-2026/}
}