Open Source vs Closed Source AI Models for Financial Analysis in 2026

Closed source models scored higher on financial analysis benchmarks in September 2026, but several open source models reach similar scores at a lower cost per task. For example in the Vals AI Finance Agent v2 benchmark, Gemini 3.8 Flash scores 61.4% with $2.00 per task while GLM 5.3 Flash, the highest scoring open weight model, scores 57.9% at $0.05 per task. This blogpost analyses results of closed and open source models on financial analysis, and sets it in relation to cost and speed.

Accuracy on finance

Finance Agent v2 ranking

Finance Agent v2 from Vals AI tests models on 927 analyst questions based on SEC filings. Each model uses six tools: EDGAR search, web search, an HTML page parser, a retrieval tool, a calculator and a price history tool. The scores below are the mean of three runs on the private test set of 450 questions, as published on 23 September 2026. Cost per task is the API cost of one question on the provider Vals AI used for the run. Time per task is the average time from question to final answer.

RankModelDeveloperTypeAccuracyCost per taskTime per task
1Gemini 3.8 FlashGoogleClosed source61.4%$2.003.3 min
2Muse Spark 1.2MetaClosed source60.6%$0.775.1 min
3Muse Spark 1.3 MaxMetaClosed source60.0%$0.763.3 min
4Gemini 3.7 FlashGoogleClosed source59.0%$1.483.0 min
5Muse Spark 1.3MetaClosed source58.9%$0.745.7 min
6Claude Fable 5.1AnthropicClosed source58.9%$8.3518.2 min
7Claude Opus 5AnthropicClosed source58.6%$5.1210.0 min
8Claude Opus 5.5AnthropicClosed source58.6%$9.0333.1 min
9Gemini 3.5 FlashGoogleClosed source57.9%$2.515.4 min
10GLM 5.3 FlashZ.aiOpen weight57.9%$0.0514.8 min
11MiMo V2.6 ProXiaomiOpen weight57.3%$0.2010.8 min
12Muse Spark 1.1MetaClosed source57.2%$0.726.9 min
13Claude Fable 5AnthropicClosed source56.3%$8.0610.2 min
14Gemini 3.6 FlashGoogleClosed source56.3%$1.403.7 min
15MiMo V2.6 FlashXiaomiOpen weight56.3%$0.077.0 min
16GLM 5.3Z.aiOpen weight55.8%$1.0715.8 min
17Hy4 PreviewTencentOpen weight55.1%$0.5923.7 min
18GPT 5.6 LunaOpenAIClosed source55.0%$0.2612.9 min
19Ling 3.0 Flash FinAnt GroupOpen weight54.9%$0.046.3 min
20GPT 5.6 TerraOpenAIClosed source54.4%$3.3624.2 min

The top 20 models include six open weight models: GLM 5.3 Flash in 10th place, MiMo V2.6 Pro in 11th, MiMo V2.6 Flash in 15th, GLM 5.3 in 16th, Hy4 Preview in 17th and Ling 3.0 Flash Fin in 19th. Kimi K3 follows in 21st place at 54.4% and DeepSeek V4.1 Flash in 27th place at 53.5%. Four of the six, MiMo V2.6 Pro, MiMo V2.6 Flash, Hy4 Preview and Ling 3.0 Flash Fin, were added to the leaderboard after the 11 September 2026 update. The full ranking of open weight models with parameters, licenses and hardware requirements is in Best open source models for financial analysis.

Open vs closed models as groups

The group figures cover 59 of the 68 models on the leaderboard: 20 open weight models, whose weights are published on Hugging Face, and 39 closed source models, which are available through an API only. The remaining 9 models are left out of the group figures because their weight availability is unconfirmed.

MeasureOpen weight modelsClosed source models
Models on the leaderboard2039
Highest accuracy57.9% (GLM 5.3 Flash)61.4% (Gemini 3.8 Flash)
Average accuracy of the top 556.5%60.0%
Median accuracy of all models in the group50.0%52.3%
Models at 50% accuracy or higher1023
Median cost per task of the top 5$0.20$0.77
Median time per task of the top 514.8 min3.3 min

The top 5 open weight models are GLM 5.3 Flash, MiMo V2.6 Pro, MiMo V2.6 Flash, GLM 5.3 and Hy4 Preview. The top 5 closed source models are Gemini 3.8 Flash, Muse Spark 1.2, Muse Spark 1.3 Max, Gemini 3.7 Flash and Muse Spark 1.3.

Excel modeling and tax benchmarks

Vals AI runs two further finance benchmarks with tools. The Excel Modeling Benchmark has 103 tasks in which the model builds LBO, DCF, M&A, operating and comparable company models in Excel. Tax Agent Bench has 391 research questions on US corporate tax.

BenchmarkHighest closed source scoreHighest open weight scoreDifferenceLast updated
Finance Agent v261.4% (Gemini 3.8 Flash)57.9% (GLM 5.3 Flash)3.6 points23 Sep 2026
Tax Agent Bench77.6% (Claude Fable 5.1)73.1% (GLM 5.3)4.5 points23 Sep 2026
Excel Modeling Benchmark76.7% (Claude Fable 5.1)66.4% (Kimi K3)10.3 points22 Sep 2026

On Tax Agent Bench, GLM 5.3 ranks third of 28 models, above Muse Spark 1.3, Grok 4.6, Claude Opus 5.5 and GPT 5.6 Sol. On the Excel Modeling Benchmark, Kimi K3 ranks 15th of 65 models, and the next open weight models are MiMo V2.6 Flash at 65.5%, MiMo V2.6 Pro at 62.9% and GLM 5.2 at 61.5%.

Results by task

Finance Agent v2 scores each model on nine task categories and on All Pass, a strict metric that counts an answer as correct only when every check in it passes. The table compares the highest closed source score with the highest open weight score in each category.

CategoryHighest closed source scoreHighest open weight scoreDifference
Disclosure Analysis71.3% (Claude Opus 5)71.5% (GLM 5.3 Flash)Open weight 0.3 points higher
General Quantitative Analysis81.8% (Gemini 3.7 Flash)79.8% (MiMo V2.6 Pro)2.0 points
Earnings Analysis81.4% (Gemini 3.8 Flash)78.7% (GLM 5.3)2.7 points
Comparables52.0% (Gemini 3.8 Flash)49.0% (MiMo V2.6 Flash)3.0 points
Adjustments56.3% (Muse Spark 1.2)53.2% (GLM 5.3 Flash)3.1 points
Overall61.4% (Gemini 3.8 Flash)57.9% (GLM 5.3 Flash)3.6 points
All Pass50.9% (Muse Spark 1.2)46.3% (GLM 5.3 Flash)4.6 points
Precedents36.4% (Gemini 3.5 Flash)30.9% (MiMo V2.6 Pro)5.5 points
General Qualitative Analysis83.4% (Muse Spark 1.3 Max)76.2% (GLM 5.3 Flash)7.1 points
Market Analysis78.5% (Gemini 3.8 Flash)70.4% (GLM 5.3 Flash)8.1 points
Financial Modeling34.5% (Muse Spark 1.2)25.4% (GLM 5.3 Flash)9.1 points

Smallest gaps

  • Disclosure Analysis: GLM 5.3 Flash scores 71.5% and Claude Opus 5 scores 71.3%. This category tracks changes in MD&A language, KPI definitions and segment reporting across several annual filings.
  • General Quantitative Analysis: MiMo V2.6 Pro scores 79.8% against 81.8% for Gemini 3.7 Flash. This category covers the extraction and calculation of reported financials such as revenue growth, CAGR and leverage ratios.
  • Earnings Analysis: GLM 5.3 scores 78.7% against 81.4% for Gemini 3.8 Flash. This category compares reported results with consensus estimates, prior guidance and non GAAP reconciliations.

Largest gaps

  • Financial Modeling: Muse Spark 1.2 scores 34.5% and GLM 5.3 Flash scores 25.4%, a difference of 9.1 points. This category covers DCF, LBO and accretion and dilution models built from filing data.
  • Market Analysis: Gemini 3.8 Flash scores 78.5% and GLM 5.3 Flash scores 70.4%. This category covers stock price reactions, total shareholder return and volatility against sector indices.
  • General Qualitative Analysis: Muse Spark 1.3 Max scores 83.4% and GLM 5.3 Flash scores 76.2%. This category covers the summary and comparison of business models, risk factors and MD&A across companies.

Cost and speed

Cost per task

The table pairs open weight models with closed source models that reach a similar score on Finance Agent v2. Costs are API prices from the Vals AI run. Running an open weight model on the firm's own GPUs replaces the per token price with the cost of the hardware and its operation.

Open weight modelAccuracyCost per taskClosed source modelAccuracyCost per task
GLM 5.3 Flash57.9%$0.05Gemini 3.5 Flash57.9%$2.51
GLM 5.3 Flash57.9%$0.05Claude Opus 558.6%$5.12
MiMo V2.6 Pro57.3%$0.20Muse Spark 1.157.2%$0.72
MiMo V2.6 Flash56.3%$0.07Gemini 3.6 Flash56.3%$1.40
Hy4 Preview55.1%$0.59GPT 5.6 Luna55.0%$0.26
Ling 3.0 Flash Fin54.9%$0.04GPT 5.6 Terra54.4%$3.36
Kimi K354.4%$1.07Claude Sonnet 553.9%$1.13
DeepSeek V4.1 Flash53.5%$0.21GPT 6 Astra53.5%$6.82

Ling 3.0 Flash Fin has the lowest cost per task of the 20 highest scoring models at $0.04, followed by GLM 5.3 Flash at $0.05 and MiMo V2.6 Flash at $0.07. GLM 5.3 Flash reaches the same score as Gemini 3.5 Flash at 2% of the cost. The cost figures depend on the provider: Vals AI ran GLM 5.3 Flash and GLM 5.3 through Fireworks AI, and the MiMo, Hy4, Ling, DeepSeek, Kimi and Qwen models through their developers' own APIs. Among closed source models, the Muse Spark models reach 57 to 61% at $0.72 to $0.77 per task, and GPT 5.6 Luna reaches 55.0% at $0.26.

Time per task

The top 5 open weight models take a median of 14.8 minutes per task on Finance Agent v2, and the top 5 closed source models take 3.3 minutes. GLM 5.3 Flash takes 14.8 minutes, MiMo V2.6 Pro takes 10.8 minutes, GLM 5.3 takes 15.8 minutes and Hy4 Preview takes 23.7 minutes. Gemini 3.8 Flash takes 3.3 minutes and Muse Spark 1.3 Max takes 3.3 minutes. MiMo V2.6 Flash takes 7.0 minutes, the shortest time among the top 5 open weight models, and Ling 3.0 Flash Fin takes 6.3 minutes. Time per task includes every tool call and depends on the inference provider and the reasoning effort setting, which Vals AI set to max for the GLM models.

Data control and licensing

CriterionOpen weight modelsClosed source models
Where the model runsOn the firm's own servers, in its own cloud account, or through a hosted providerOn the developer's servers, accessed through an API
Where prompts and documents goStay inside the firm's infrastructure when self hostedSent to the developer under its API data terms
Model versionFixed until the firm changes itUpdated or retired on the developer's schedule
Fine tuning on internal dataPossible on the published weightsLimited to the options the developer offers
Cost structureHardware and operations when self hosted, or per token through a providerPer token or per request
LicenseMIT, Apache 2.0 or a custom license with conditions (see the license table in the open source models ranking)Developer's commercial terms of service
Accuracy on Finance Agent v2Highest score 57.9%Highest score 61.4%

FAQ

Are closed source models better than open source models for financial analysis?

Closed source models have the higher scores on the three Vals AI finance benchmarks with tools as of September 2026. The difference is 3.6 points on Finance Agent v2, 4.5 points on Tax Agent Bench and 10.3 points on the Excel Modeling Benchmark. On the Disclosure Analysis category of Finance Agent v2, GLM 5.3 Flash scores 0.3 points above the highest closed source model.

What is the best open source model compared with GPT, Claude and Gemini for finance?

GLM 5.3 Flash from Z.ai has the highest open weight score on Finance Agent v2 at 57.9%, followed by MiMo V2.6 Pro from Xiaomi at 57.3%. Gemini 3.8 Flash scores 61.4%, Claude Opus 5 scores 58.6% and GPT 5.6 Luna scores 55.0%. GLM 5.3 Flash ranks 10th of 68 models and costs $0.05 per task.

Are open source models cheaper than closed source models for financial analysis?

At similar accuracy, several open weight models cost less per task through hosted APIs. GLM 5.3 Flash costs $0.05 per task at 57.9%, and Claude Opus 5 costs $5.12 per task at 58.6%. The Muse Spark models reach 57 to 61% at about $0.75 per task. The median cost per task of the top 5 models is $0.20 for open weight models and $0.77 for closed source models.

Why do open source models take longer per task?

On Finance Agent v2, the top 5 open weight models take a median of 14.8 minutes per task and the top 5 closed source models take 3.3 minutes. Time per task depends on the number of tool calls, the reasoning effort setting and the speed of the inference provider. Vals AI ran the GLM models at max reasoning effort through Fireworks AI.

Which financial tasks show the largest difference between open and closed source models?

Financial Modeling shows the largest difference on Finance Agent v2 at 9.1 points, followed by Market Analysis at 8.1 points and General Qualitative Analysis at 7.1 points. On the Excel Modeling Benchmark, which tests model building in Excel, the difference is 10.3 points.

Last updated: 25 September 2026. Benchmark scores, prices and model versions change with each release. Educational content for general information about AI models.