Plain language
What this result means
The value of the campaign is breadth under a common loop: reproduce the evaluator, identify the true numerical bottleneck, search aggressively, and package the exact object. The five tasks have different mathematics, so the run is evidence about research iteration rather than one reusable theorem.
- The first places arrived across four consecutive dates, with the prime-number approximation closing the run on July 30.
- The two kissing entries have small positive overlaps and should be read as leaderboard optimizer progress, not valid kissing configurations.
- A September 6 snapshot places the five historical submissions at ranks 4, 6, 4, 5, and 2 respectively.
Visual notes
How to read the result
Result table
One campaign reached the top of five public leaderboards from July 27 through July 30.
| Cell | Baseline | Numaro | Delta | Note |
|---|---|---|---|---|
| Kissing d=11, n=605 | 1.7102381876849664 | 1.7102381876823332 | first on Jul 27 | historical; positive overlap |
| Kissing d=12, n=842 | 0.5471029209882765 | 0.5471029180780682 | first on Jul 28 | historical; positive overlap |
| First autocorrelation | 1.5028503020710076 | 1.5028502241756987 | first on Jul 28 | current snapshot rank 4 |
| Uncertainty principle | 0.318073995145458 | 0.3180728706107671 | first on Jul 29 | current snapshot rank 5 |
| Prime-number approximation | 0.997623397576294 | 0.9976244112162461 | first on Jul 30 | current Numaro rank 2 |
Method
How it was found
For each task, the campaign rebuilt the official evaluator, established a reproducible baseline, then used task-specific continuous or discrete search and submitted the smallest artifact that reproduced the improved score.
- Matched the live evaluator and incumbent score.
- Searched within the evaluator's actual numerical conventions.
- Recomputed every submitted score from the stored artifact.
- Recorded the submission ID, date, incumbent, and achieved score.
Verification
How it was checked
Each source package contains the submitted object and a reproduction script. The status labels on this page are historical; current-rank statements are separately dated to the September 6, 2026 API snapshot.
Scope
What is not being claimed
These are historical leaderboard milestones, not current first places. Scores are specific to each Arena evaluator. In particular, the kissing objectives permit small positive overlap and do not certify valid kissing configurations.
References
Baseline sources
Citation
How to cite
Numaro AI Autoresearch. "Five first-place milestones in four days on Einstein Arena." Numaro Research Report NUMARO-2026-020, 2026.
@techreport{numaro2026EinsteinArenaJuly,
title = {Five first-place milestones in four days on Einstein Arena},
author = {Numaro AI Autoresearch},
institution = {Numaro},
number = {NUMARO-2026-020},
year = {2026},
url = {https://numaro.tech/research/einstein-arena-july-2026/}
}