Research

Research on evaluating and improving large language models for institutional financial analysis: public leaderboards, controlled studies, and reproducible harnesses. Scores are published; proprietary inputs must be private.

The Aiera Leaderboard

A pure-research leaderboard for financial LLMs: how well does a model perform the analytical work of a research analyst, grounded in professional data. Every model is scored on three equal axes: Facts (right and complete), Grounding (cited and defensible), and Depth (institutional-caliber insight). The headline finding: factual completeness and analytical depth diverge at the top; some models that lead on accuracy produce thorough but shallow analysis.

3 axes
Facts, Grounding, and Depth; equally weighted

The Aiera Lift

How much does connecting any model to Aiera's MCP server improve research answering quality and depth? Large gains on the ~78% of questions the open web can't reach, but only for models that can drive the tools. The discriminator is tool-use competence, rather than capability tier: some open-weight models lift while a frontier model doesn't.

+0.30
Best model's key-facts gain on proprietary questions; ~2.4× the field's average lift

Corpus-Derived Style Templates

Can the register of professional sell-side research be distilled automatically from a large corpus and then injected as a system-prompt style template? Yes, a single global template beats a prose-capable no-template baseline decisively, the gains are substantive (key-fact recall rises), and per-topic specialization isn't even required, even under oracle routing.

90.5%
Preferred over a no-template baseline, blinded pairwise, strong model preference