Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
Creating Post-Training Datasets for Research-Level Mathematics (matharena.ai)
3 points by frozenseven 13 days ago | past | discuss
ArXivLean: How Well Can LLMs Formally Prove Research Math? (matharena.ai)
3 points by OxfordCommand 4 months ago | past
BrokenArXiv: How often do LLMs claim to prove false theorems? (matharena.ai)
3 points by robinhouston 5 months ago | past
MathArena: Evaluating LLMs on uncontaminated math questions (matharena.ai)
2 points by GaggiX 6 months ago | past
New open source model achieves same score as GPT 5.2 High on AIME2026 I (matharena.ai)
3 points by mh3467 7 months ago | past | 2 comments
MathArena Apex: Unconquered Final-Answer Problems (matharena.ai)
2 points by frozenseven 11 months ago | past
Evaluating publicly available LLMs on IMO 2025 (matharena.ai)
79 points by hardmaru on July 19, 2025 | past | 89 comments
Not Even Bronze: Evaluating LLMs on 2025 International Math Olympiad (matharena.ai)
3 points by amichail on July 19, 2025 | past | 2 comments
IMO 2025 LLM results are in (matharena.ai)
5 points by arberavdullahu on July 18, 2025 | past | 1 comment
Not Even Bronze? Evaluating LLMs on 2025 International Math Olympiad (matharena.ai)
1 point by EvgeniyZh on July 18, 2025 | past | 1 comment
Gemini 2.5 gets 24.4% on MathArena USAMO beating previous top score of 4.7% (matharena.ai)
54 points by alphabetting on April 2, 2025 | past | 10 comments
OpenAI o3-mini scores 78% on yesterday's AIME 2025 math competition (matharena.ai)
3 points by bmislav on Feb 7, 2025 | past

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: