Finance research agents face a fundamental evaluation problem: rubrics must reflect expert standards while remaining fixed relative to an information cutoff date. Traditional benchmarks rely on hand-crafted, per-item rubrics that are expensive to produce and cannot capture the varying standards across institutions or research domains. FinAutoRubric addresses this by separating expert guidance from query-specific rubric generation, letting reusable criteria from a Task Bank carry across tasks while agents and code produce the actual rubric for each evaluation query.

Problem with Fixed Rubrics and the Need for Adaptivity

In typical finance benchmarks, each evaluation item carries a rigid rubric defined by human experts. These rubrics encode what a "correct" answer looks like, but they are brittle: adding a new task often requires defining an entirely new rubric from scratch, and different institutions may want different criteria or weighting. Moreover, as language models improve, rubrics designed for earlier models quickly become outdated, leaving unclear how well newer agents perform relative to the same standard. The paper identifies this as a scaling problem—without a systematic way to generate and update rubrics, evaluation either stagnates or becomes disproportionately costly as the number of tasks and models grows.

FinAutoRubric Architecture

The core of FinAutoRubric is a two-level design. At the top, experts provide reusable evaluation guidance in the form of high-level criteria, principles, and expected answer patterns. This guidance is encoded as prompts for language model agents and as syntactic rules that validating code enforces. Below this, a Task Bank stores these reusable criteria, indexed by task type, asset class, or research objective. When an evaluation query arrives, two agents operate in sequence: a writer agent researches and drafts a rubric for the specific query, guided by the expert prompts and rules; a reviewer agent then verifies the drafted rubric against the expert guidance, flagging inconsistencies or gaps. If the reviewer cannot resolve flagged issues automatically, the failure escalates to a human analyst. This loop operates over multiple iterations until the rubric meets expert validation or a human intervenes.

The writer agent does not generate rubrics from scratch; it maps the query onto the Task Bank, retrieves relevant criteria, and expands them into a full rubric structure. The reviewer agent checks for completeness, ensuring that every expected value or reasoning step required by the expert guidance is present, and that no extraneous or contradictory elements have been introduced. Because both agents operate under the same expert guidance—the same prompts and the same code-enforced rules—the resulting rubrics are coherent across tasks and models.

Query-Specific Rubric Generation and Validation

The paper emphasizes that rubric generation is not a one-shot process. In long-horizon loops that follow the expert guidance, the writer researches every expected value relevant to the query, and the reviewer verifies each one. This back-and-forth continues until the rubric either passes validation or enough iterations have occurred to trigger human escalation. The expert guidance thus serves as both a prompt for the agents and a contract that the generated rubric must satisfy. The code that enforces the rules is deterministic: if the generated rubric violates a structural constraint (e.g., missing a required criterion category), it is automatically rejected and the writer regenerates. This reduces the burden on human experts, who only need to review cases where the automated loop fails.

Experimental Results Across Finance Benchmarks

The authors evaluate FinAutoRubric on three expert-authored finance benchmarks. On each, the generated rubrics track expert scoring as closely as the strongest previously evaluated generator, while stating the expert rubric's expected value for a greater number of criteria. Human grading agreement is strong: the rubrics' scores agree with human-graded answers at a level comparable to top-tier generator rubrics. In a blind review, in-house analysts preferred FinAutoRubric-generated rubrics over baselines, citing clarity and consistency as key factors. The authors also released a 100-query FinAutoRubric Benchmark, constructed from in-house analysts' key questions across 78 tasks and eight asset classes. A notable finding is that rubrics generated by earlier model versions still leave headroom for later models—indicating that the benchmark and rubric framework capture enduring evaluation criteria rather than overfitting to a single model generation.

Limitations and Trade-offs

The automated rubric generation loop is not without overhead. The writer-reviewer cycle, especially when it escalates to human analysts, can be time-consuming for complex or unusual queries. The quality of the generated rubric depends on the completeness of the expert guidance and the Task Bank; if the expert principles are underspecified, the agents may generate rubrics that miss important subtleties. Additionally, the benchmark was built from in-house analysts' questions, which may not capture the full diversity of institutional standards or emerging asset classes. The authors note that extending the Task Bank and expert guidance to new domains requires careful curation, and that the current system is most effective when the expert guidance is already well-structured.

Practical Impact for Developers and Researchers

For developers working on financial AI agents, FinAutoRubric provides a practical method for keeping evaluation aligned with expert standards without redefining rubrics from scratch each time a new model is released. The separation of expert guidance from query-specific generation means that adding a new task primarily involves indexing relevant criteria from the Task Bank, not writing a new rubric from scratch. The blind-review preference from in-house analysts suggests that the generated rubrics are not only accurate but also more interpretable and easier to use than many hand-crafted alternatives. Researchers can also use the released 100-query benchmark to evaluate their own agents against a standardized set of finance questions, with the assurance that the rubrics evolve alongside model capabilities.

Conclusion

FinAutoRubric replaces fixed, per-item rubrics with an expert-guided, agent-driven generation pipeline that can adapt to new tasks and models while preserving the consistency that expert oversight provides. By combining a Task Bank of reusable criteria with writer and reviewer agents operating under deterministic code-enforced rules, the system produces rubrics that track expert scoring closely, agree with human grading, and are preferred in blind reviews. The released benchmark and open-code framework offer an immediate resource for anyone evaluating or building financial research agents, and the architectural pattern—expert guidance coupled with query-specific generation—may generalize to other domains where evaluation criteria must remain stable yet adaptable.

Read the paper on arXiv