Research on evaluating and improving large language models for institutional financial analysis: public leaderboards, controlled studies, and reproducible harnesses. Scores are published; proprietary inputs must be private.
A pure-research leaderboard for financial LLMs: how well does a model perform the analytical work of a research analyst, grounded in professional data. Every model is scored on three equal axes: Facts (right and complete), Grounding (cited and defensible), and Depth (institutional-caliber insight). The headline finding: factual completeness and analytical depth diverge at the top; some models that lead on accuracy produce thorough but shallow analysis.
How much does connecting any model to Aiera's MCP server improve research answering quality and depth? Large gains on the ~78% of questions the open web can't reach, but only for models that can drive the tools. The discriminator is tool-use competence, rather than capability tier: some open-weight models lift while a frontier model doesn't.
Can the register of professional sell-side research be distilled automatically from a large corpus and then injected as a system-prompt style template? Yes, a single global template beats a prose-capable no-template baseline decisively, the gains are substantive (key-fact recall rises), and per-topic specialization isn't even required, even under oracle routing.