This finance-model benchmark card is more useful for what it discloses than for who "wins"
The official benchmark card for Ling-3.0-flash-Fin is a useful reminder that the unit being tested is rarely just “the model.” The release says most runs used temperature 1, top_p 0.95 and the highest available reasoning effort. FinFIRST and FinSearchComp Verified used a common ReAct scaffold with Web Search, Visit and Python. SpreadsheetBench used Claude Code 2.1.173 with LibreOffice 25.8.7, Search disabled, 120 or 300 maximum interaction turns and a three-hour task timeout; Ling used temperature 0.6 there. The chart also mixes evidence types. Some results come from official or externally published scores, while others are internal runs. FinSearchComp Verified is an internal 145-question set with expert-revised answers and a GPT-5 judge. FinCRAFT is internal. FinFIRST is announced as “coming soon,” not public today. None of that makes the chart useless. It makes the claim narrower: these are reported results under several specific agent systems, tool budgets and evaluation pipelines—not a clean intrinsic ranking of raw checkpoints. The finance weights are also not public yet; the team says they are due next week. Once they land, the most valuable follow-up would be the exact harnesses, prompts, tool adapters, per-run variance and failure traces. Until then, the bars are a test plan, not an independent reproduction.