From research to an algorithm
Relaxation proposes. Exact scoring decides.
A smooth model gives the search a direction. A legal move is kept only when the exact benchmark score improves. After discovery, the placer runs entirely as GPU code.
Make the grid navigable.
Spread each pin's demand across neighboring cells without changing its total mass. Use gradients to propose global movement.
Approximate guidanceKeep verified improvements.
Check legality and the exact score. Reject a misleading proposal; retain the incumbent. Incremental GPU scoring makes large batches practical.
Exact acceptanceComparison with published results. Archgen remains the official challenge winner. Numaro's score is reverified by the campaign's standalone verifier and has not been confirmed by the organizers.
Inspect all 17 circuits 900 s · seed 0 / 300 s · three seeds
| Circuit | Split | 900 s | 300 s mean | Legal |
|---|---|---|---|---|
| ibm01 | Development | 0.75559 | 0.76578 | Yes |
| ibm02 | Validation | 0.91330 | 0.92367 | Yes |
| ibm03 | Held out | 0.85983 | 0.86747 | Yes |
| ibm04 | Development | 0.88570 | 0.89643 | Yes |
| ibm06 | Development | 0.98019 | 1.00380 | Yes |
| ibm07 | Development | 0.94164 | 0.96285 | Yes |
| ibm08 | Held out | 0.98316 | 0.99398 | Yes |
| ibm09 | Development | 0.75233 | 0.75752 | Yes |
| ibm10 | Validation | 0.89921 | 0.91127 | Yes |
| ibm11 | Development | 0.76116 | 0.77410 | Yes |
| ibm12 | Development | 1.06622 | 1.07678 | Yes |
| ibm13 | Held out | 0.80309 | 0.82143 | Yes |
| ibm14 | Development | 1.03093 | 1.03845 | Yes |
| ibm15 | Development | 0.98955 | 1.01847 | Yes |
| ibm16 | Validation | 0.95798 | 0.97174 | Yes |
| ibm17 | Development | 1.16362 | 1.16364 | Yes |
| ibm18 | Held out | 1.17915 | 1.16305 | Yes |
| Mean | All circuits | 0.93663 | 0.94767 | All |
The frozen method commit is f264d1f8. Held-out circuits are ibm03, ibm08, ibm13, and ibm18.
What happened after routing? Six placements of ariane136
All six placements routed successfully with the same RTL and flow. Lower proxy correlated with lower power; the wirelength trend was not statistically significant, and timing had no relationship with proxy. Total negative slack was zero in all six runs.
A separate exploratory comparison with OpenROAD RTL-MP improved some routing-demand and cell-density tails but worsened mean congestion and IR drop. The two placements needed different macro halos, so that comparison was not controlled. An earlier timing-improvement claim was retracted after the six-placement study.
Plain language
What this result means
Before a chip's small logic cells are placed and its wires routed, large memory blocks and hard IP need locations on the floorplan. Those choices shape the space available to every later tool. Here the AI performed the research that created a placement algorithm: it proposed hypotheses, wrote code, ran experiments, and recorded what survived. The delivered algorithm uses no language model or trained placement model at run time.
- The 900-second run lowers the published rank-1 mean from 0.9507 to 0.93663, a 1.48% reduction. The comparison uses reported baseline scores on hardware we do not control.
- All three 300-second seed means—0.94808, 0.94690, and 0.94803—are below 0.9507. These are complete 17-circuit sweeps, not a selection of the best circuits.
- Ten development circuits, three validation circuits, and four test circuits were assigned before optimization. The algorithm and parameters were frozen before test evaluation.
- Twelve search modifications were rejected. The useful changes made the congestion objective smoother for guidance and exact evaluation cheaper for decisions.
Visual notes
How to read the result

Result table
A lower mean proxy cost across all 17 benchmark circuits.
| Cell | Baseline | Numaro | Delta | Note |
|---|---|---|---|---|
| Archgen, published rank 1 | 0.9507 | 0.93663 | −1.48% | 900 s/circuit · seed 0 · one L40S |
| Carrotato / AbuPlace, rank 2 | 0.9522 | 0.93663 | −1.64% | published baseline mean |
| JaneRT, rank 3 | 0.9694 | 0.93663 | −3.38% | published baseline mean |
| RePlAce | 1.4578 | 0.93663 | −35.75% | same 17-circuit suite |
| 300-second search budget | 0.9507 | 0.94767 | −0.32% | mean across 3 seeds · all 51 placements legal |
Method
How it was found
The campaign ran for 48 hours on one L40S. A human chose the problem and compute budget and reviewed the result; the research system generated hypotheses, implemented experiments, and assigned verdicts. The method jointly moves hard macros and cell clusters, first using a differentiable relaxation and then batched discrete search. Both stages keep only exact-scored improvements.
- Start from the benchmark placement and legalize the hard macros.
- Replace pin snapping with bilinear mass distribution and route snapping with continuous cell overlap, preserving total demand.
- Use the relaxed objective to propose global moves, then score legal checkpoints with the exact benchmark objective.
- Batch discrete proposals on the GPU. Incremental scoring updates the moved macro's affected nets; combined moves are rescored before acceptance.
- Return the best exact-scored legal placement, then verify its coordinates independently of the search implementation.
Verification
How it was checked
The campaign's standalone verifier imports nothing from the search code. It reloads the original benchmark netlists through the reference TILOS evaluator, recomputes scores from the final coordinates, and checks overlaps, bounds, and fixed macros. The campaign recorded component agreement within 10⁻¹⁴ relative error over 68 full-evaluator comparisons and maximum incremental deviation of 2.2×10⁻¹⁶. For this page, the aggregate and seed means were recalculated from the stored result records; the original GPU campaign and OpenROAD flows were not rerun.
Scope
What is not being claimed
This is a lower reported benchmark proxy, not an official challenge win, an optimality proof, or a general claim of better chips. The four test circuits belong to the same benchmark family. Physical validation is limited to two designs, with a controlled six-placement study on one design whose timing constraint does not bind. Only the power correlation was significant. An exploratory comparison with OpenROAD RTL-MP improves some tail metrics but worsens mean congestion and IR drop, and uses different macro halos. Smooth optimization, exact acceptance, and incremental evaluation also appear in prior work; the paper isolates its differences from AbuPlace. LLM token usage and cost were not instrumented.
References
Baseline sources
- Research manuscript (preprint, PDF).
- Frozen placer, verifier, placements, and campaign evidence (ZIP).
- Per-circuit scores, seed statistics, and provenance (JSON).
- Partcl/HRT challenge: public benchmark, rules, and official leaderboard.
- AbuPlace: the closest algorithmic comparator.
- Archgen: the published challenge-winning implementation.
Citation
How to cite
Numaro AI Autoresearch. "AI Autoresearch for Chip Design: Autonomous Discovery of a Macro Placement Algorithm." Numaro Research Report NUMARO-2026-022, 2026.
@techreport{numaro2026MacroPlacementRelax,
title = {AI Autoresearch for Chip Design: Autonomous Discovery of a Macro Placement Algorithm},
author = {Numaro AI Autoresearch},
institution = {Numaro},
number = {NUMARO-2026-022},
year = {2026},
url = {https://numaro.tech/research/macro-placement-relax-and-gate-2026/}
}