IRAB: A Benchmark for Investment Research Agents on Real-World Tasks
Initial Technical Report
Investment research is a high-stakes, data-intensive workflow: what matters is gathering the right numbers from the right primary documents and reasoning over them correctly, not how the write-up reads. Existing finance benchmarks, however, tend to score prose-style answers, freeze their gold answers for reproducibility, which invites contamination, and rarely check that their own grader can be trusted. We present IRAB 1.0, the Investment Research Agent Benchmark, an evaluation system for LLM agents built to close this gap. Its tasks are randomly sampled from real investment-research usage and then de-identified and rewritten for release; its gold answers carry a shelf life—kept live and refreshed rather than frozen, so the benchmark stays current and contamination-resistant; and its agentic harness ships a built-in data pipeline, so agents are measured on the data-acquisition work the job actually requires. Each task is graded by rubrics mined from real practitioner pain points and aggregated as Final = Data × Quality × Reliability—a data gate under which a fluent answer built on the wrong data scores near zero. On a preliminary first round over 101 held-out tasks and nine frontier models (eleven configurations including reasoning-on/off variants), GPT 5.5 leads at a Final score of 0.66, but the field is closely spaced and far from solved (Figure 1). We further verify that the grading judge is stable across reruns (r ≈ 0.91), consistent with expert review on 82% of scores, and free of measurable self-preference, which gives us some confidence in the comparison at this early stage.
Contents
- Introduction
- The IRAB Benchmark
- Evaluation Framework
- Experiments and Results
- Judge Stability and Expert Consistency
- Future work
- Conclusion and Limitations
Live leaderboard · Benchmark repository