Open Source vs Closed Source AI Models for Financial Analysis in 2026
Closed source models scored higher on financial analysis benchmarks in September 2026, but several open source models reach similar scores at a lower cost per task. For example in the Vals AI Finance Agent v2 benchmark, Gemini 3.8 Flash scores 61.4% with $2.00 per task while GLM 5.3 Flash, the highest scoring open weight model, scores 57.9% at $0.05 per task. This blogpost analyses results of closed and open source models on financial analysis, and sets it in relation to cost and speed.
On this page:
Accuracy on finance
Finance Agent v2 ranking
Finance Agent v2 from Vals AI tests models on 927 analyst questions based on SEC filings. Each model uses six tools: EDGAR search, web search, an HTML page parser, a retrieval tool, a calculator and a price history tool. The scores below are the mean of three runs on the private test set of 450 questions, as published on 23 September 2026. Cost per task is the API cost of one question on the provider Vals AI used for the run. Time per task is the average time from question to final answer.
| Rank | Model | Developer | Type | Accuracy | Cost per task | Time per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash | Closed source | 61.4% | $2.00 | 3.3 min | |
| 2 | Muse Spark 1.2 | Meta | Closed source | 60.6% | $0.77 | 5.1 min |
| 3 | Muse Spark 1.3 Max | Meta | Closed source | 60.0% | $0.76 | 3.3 min |
| 4 | Gemini 3.7 Flash | Closed source | 59.0% | $1.48 | 3.0 min | |
| 5 | Muse Spark 1.3 | Meta | Closed source | 58.9% | $0.74 | 5.7 min |
| 6 | Claude Fable 5.1 | Anthropic | Closed source | 58.9% | $8.35 | 18.2 min |
| 7 | Claude Opus 5 | Anthropic | Closed source | 58.6% | $5.12 | 10.0 min |
| 8 | Claude Opus 5.5 | Anthropic | Closed source | 58.6% | $9.03 | 33.1 min |
| 9 | Gemini 3.5 Flash | Closed source | 57.9% | $2.51 | 5.4 min | |
| 10 | GLM 5.3 Flash | Z.ai | Open weight | 57.9% | $0.05 | 14.8 min |
| 11 | MiMo V2.6 Pro | Xiaomi | Open weight | 57.3% | $0.20 | 10.8 min |
| 12 | Muse Spark 1.1 | Meta | Closed source | 57.2% | $0.72 | 6.9 min |
| 13 | Claude Fable 5 | Anthropic | Closed source | 56.3% | $8.06 | 10.2 min |
| 14 | Gemini 3.6 Flash | Closed source | 56.3% | $1.40 | 3.7 min | |
| 15 | MiMo V2.6 Flash | Xiaomi | Open weight | 56.3% | $0.07 | 7.0 min |
| 16 | GLM 5.3 | Z.ai | Open weight | 55.8% | $1.07 | 15.8 min |
| 17 | Hy4 Preview | Tencent | Open weight | 55.1% | $0.59 | 23.7 min |
| 18 | GPT 5.6 Luna | OpenAI | Closed source | 55.0% | $0.26 | 12.9 min |
| 19 | Ling 3.0 Flash Fin | Ant Group | Open weight | 54.9% | $0.04 | 6.3 min |
| 20 | GPT 5.6 Terra | OpenAI | Closed source | 54.4% | $3.36 | 24.2 min |
The top 20 models include six open weight models: GLM 5.3 Flash in 10th place, MiMo V2.6 Pro in 11th, MiMo V2.6 Flash in 15th, GLM 5.3 in 16th, Hy4 Preview in 17th and Ling 3.0 Flash Fin in 19th. Kimi K3 follows in 21st place at 54.4% and DeepSeek V4.1 Flash in 27th place at 53.5%. Four of the six, MiMo V2.6 Pro, MiMo V2.6 Flash, Hy4 Preview and Ling 3.0 Flash Fin, were added to the leaderboard after the 11 September 2026 update. The full ranking of open weight models with parameters, licenses and hardware requirements is in Best open source models for financial analysis.
Open vs closed models as groups
The group figures cover 59 of the 68 models on the leaderboard: 20 open weight models, whose weights are published on Hugging Face, and 39 closed source models, which are available through an API only. The remaining 9 models are left out of the group figures because their weight availability is unconfirmed.
| Measure | Open weight models | Closed source models |
|---|---|---|
| Models on the leaderboard | 20 | 39 |
| Highest accuracy | 57.9% (GLM 5.3 Flash) | 61.4% (Gemini 3.8 Flash) |
| Average accuracy of the top 5 | 56.5% | 60.0% |
| Median accuracy of all models in the group | 50.0% | 52.3% |
| Models at 50% accuracy or higher | 10 | 23 |
| Median cost per task of the top 5 | $0.20 | $0.77 |
| Median time per task of the top 5 | 14.8 min | 3.3 min |
The top 5 open weight models are GLM 5.3 Flash, MiMo V2.6 Pro, MiMo V2.6 Flash, GLM 5.3 and Hy4 Preview. The top 5 closed source models are Gemini 3.8 Flash, Muse Spark 1.2, Muse Spark 1.3 Max, Gemini 3.7 Flash and Muse Spark 1.3.
Excel modeling and tax benchmarks
Vals AI runs two further finance benchmarks with tools. The Excel Modeling Benchmark has 103 tasks in which the model builds LBO, DCF, M&A, operating and comparable company models in Excel. Tax Agent Bench has 391 research questions on US corporate tax.
| Benchmark | Highest closed source score | Highest open weight score | Difference | Last updated |
|---|---|---|---|---|
| Finance Agent v2 | 61.4% (Gemini 3.8 Flash) | 57.9% (GLM 5.3 Flash) | 3.6 points | 23 Sep 2026 |
| Tax Agent Bench | 77.6% (Claude Fable 5.1) | 73.1% (GLM 5.3) | 4.5 points | 23 Sep 2026 |
| Excel Modeling Benchmark | 76.7% (Claude Fable 5.1) | 66.4% (Kimi K3) | 10.3 points | 22 Sep 2026 |
On Tax Agent Bench, GLM 5.3 ranks third of 28 models, above Muse Spark 1.3, Grok 4.6, Claude Opus 5.5 and GPT 5.6 Sol. On the Excel Modeling Benchmark, Kimi K3 ranks 15th of 65 models, and the next open weight models are MiMo V2.6 Flash at 65.5%, MiMo V2.6 Pro at 62.9% and GLM 5.2 at 61.5%.
Results by task
Finance Agent v2 scores each model on nine task categories and on All Pass, a strict metric that counts an answer as correct only when every check in it passes. The table compares the highest closed source score with the highest open weight score in each category.
| Category | Highest closed source score | Highest open weight score | Difference |
|---|---|---|---|
| Disclosure Analysis | 71.3% (Claude Opus 5) | 71.5% (GLM 5.3 Flash) | Open weight 0.3 points higher |
| General Quantitative Analysis | 81.8% (Gemini 3.7 Flash) | 79.8% (MiMo V2.6 Pro) | 2.0 points |
| Earnings Analysis | 81.4% (Gemini 3.8 Flash) | 78.7% (GLM 5.3) | 2.7 points |
| Comparables | 52.0% (Gemini 3.8 Flash) | 49.0% (MiMo V2.6 Flash) | 3.0 points |
| Adjustments | 56.3% (Muse Spark 1.2) | 53.2% (GLM 5.3 Flash) | 3.1 points |
| Overall | 61.4% (Gemini 3.8 Flash) | 57.9% (GLM 5.3 Flash) | 3.6 points |
| All Pass | 50.9% (Muse Spark 1.2) | 46.3% (GLM 5.3 Flash) | 4.6 points |
| Precedents | 36.4% (Gemini 3.5 Flash) | 30.9% (MiMo V2.6 Pro) | 5.5 points |
| General Qualitative Analysis | 83.4% (Muse Spark 1.3 Max) | 76.2% (GLM 5.3 Flash) | 7.1 points |
| Market Analysis | 78.5% (Gemini 3.8 Flash) | 70.4% (GLM 5.3 Flash) | 8.1 points |
| Financial Modeling | 34.5% (Muse Spark 1.2) | 25.4% (GLM 5.3 Flash) | 9.1 points |
Smallest gaps
- Disclosure Analysis: GLM 5.3 Flash scores 71.5% and Claude Opus 5 scores 71.3%. This category tracks changes in MD&A language, KPI definitions and segment reporting across several annual filings.
- General Quantitative Analysis: MiMo V2.6 Pro scores 79.8% against 81.8% for Gemini 3.7 Flash. This category covers the extraction and calculation of reported financials such as revenue growth, CAGR and leverage ratios.
- Earnings Analysis: GLM 5.3 scores 78.7% against 81.4% for Gemini 3.8 Flash. This category compares reported results with consensus estimates, prior guidance and non GAAP reconciliations.
Largest gaps
- Financial Modeling: Muse Spark 1.2 scores 34.5% and GLM 5.3 Flash scores 25.4%, a difference of 9.1 points. This category covers DCF, LBO and accretion and dilution models built from filing data.
- Market Analysis: Gemini 3.8 Flash scores 78.5% and GLM 5.3 Flash scores 70.4%. This category covers stock price reactions, total shareholder return and volatility against sector indices.
- General Qualitative Analysis: Muse Spark 1.3 Max scores 83.4% and GLM 5.3 Flash scores 76.2%. This category covers the summary and comparison of business models, risk factors and MD&A across companies.
Cost and speed
Cost per task
The table pairs open weight models with closed source models that reach a similar score on Finance Agent v2. Costs are API prices from the Vals AI run. Running an open weight model on the firm's own GPUs replaces the per token price with the cost of the hardware and its operation.
| Open weight model | Accuracy | Cost per task | Closed source model | Accuracy | Cost per task |
|---|---|---|---|---|---|
| GLM 5.3 Flash | 57.9% | $0.05 | Gemini 3.5 Flash | 57.9% | $2.51 |
| GLM 5.3 Flash | 57.9% | $0.05 | Claude Opus 5 | 58.6% | $5.12 |
| MiMo V2.6 Pro | 57.3% | $0.20 | Muse Spark 1.1 | 57.2% | $0.72 |
| MiMo V2.6 Flash | 56.3% | $0.07 | Gemini 3.6 Flash | 56.3% | $1.40 |
| Hy4 Preview | 55.1% | $0.59 | GPT 5.6 Luna | 55.0% | $0.26 |
| Ling 3.0 Flash Fin | 54.9% | $0.04 | GPT 5.6 Terra | 54.4% | $3.36 |
| Kimi K3 | 54.4% | $1.07 | Claude Sonnet 5 | 53.9% | $1.13 |
| DeepSeek V4.1 Flash | 53.5% | $0.21 | GPT 6 Astra | 53.5% | $6.82 |
Ling 3.0 Flash Fin has the lowest cost per task of the 20 highest scoring models at $0.04, followed by GLM 5.3 Flash at $0.05 and MiMo V2.6 Flash at $0.07. GLM 5.3 Flash reaches the same score as Gemini 3.5 Flash at 2% of the cost. The cost figures depend on the provider: Vals AI ran GLM 5.3 Flash and GLM 5.3 through Fireworks AI, and the MiMo, Hy4, Ling, DeepSeek, Kimi and Qwen models through their developers' own APIs. Among closed source models, the Muse Spark models reach 57 to 61% at $0.72 to $0.77 per task, and GPT 5.6 Luna reaches 55.0% at $0.26.
Time per task
The top 5 open weight models take a median of 14.8 minutes per task on Finance Agent v2, and the top 5 closed source models take 3.3 minutes. GLM 5.3 Flash takes 14.8 minutes, MiMo V2.6 Pro takes 10.8 minutes, GLM 5.3 takes 15.8 minutes and Hy4 Preview takes 23.7 minutes. Gemini 3.8 Flash takes 3.3 minutes and Muse Spark 1.3 Max takes 3.3 minutes. MiMo V2.6 Flash takes 7.0 minutes, the shortest time among the top 5 open weight models, and Ling 3.0 Flash Fin takes 6.3 minutes. Time per task includes every tool call and depends on the inference provider and the reasoning effort setting, which Vals AI set to max for the GLM models.
Data control and licensing
| Criterion | Open weight models | Closed source models |
|---|---|---|
| Where the model runs | On the firm's own servers, in its own cloud account, or through a hosted provider | On the developer's servers, accessed through an API |
| Where prompts and documents go | Stay inside the firm's infrastructure when self hosted | Sent to the developer under its API data terms |
| Model version | Fixed until the firm changes it | Updated or retired on the developer's schedule |
| Fine tuning on internal data | Possible on the published weights | Limited to the options the developer offers |
| Cost structure | Hardware and operations when self hosted, or per token through a provider | Per token or per request |
| License | MIT, Apache 2.0 or a custom license with conditions (see the license table in the open source models ranking) | Developer's commercial terms of service |
| Accuracy on Finance Agent v2 | Highest score 57.9% | Highest score 61.4% |
FAQ
Are closed source models better than open source models for financial analysis?
Closed source models have the higher scores on the three Vals AI finance benchmarks with tools as of September 2026. The difference is 3.6 points on Finance Agent v2, 4.5 points on Tax Agent Bench and 10.3 points on the Excel Modeling Benchmark. On the Disclosure Analysis category of Finance Agent v2, GLM 5.3 Flash scores 0.3 points above the highest closed source model.
What is the best open source model compared with GPT, Claude and Gemini for finance?
GLM 5.3 Flash from Z.ai has the highest open weight score on Finance Agent v2 at 57.9%, followed by MiMo V2.6 Pro from Xiaomi at 57.3%. Gemini 3.8 Flash scores 61.4%, Claude Opus 5 scores 58.6% and GPT 5.6 Luna scores 55.0%. GLM 5.3 Flash ranks 10th of 68 models and costs $0.05 per task.
Are open source models cheaper than closed source models for financial analysis?
At similar accuracy, several open weight models cost less per task through hosted APIs. GLM 5.3 Flash costs $0.05 per task at 57.9%, and Claude Opus 5 costs $5.12 per task at 58.6%. The Muse Spark models reach 57 to 61% at about $0.75 per task. The median cost per task of the top 5 models is $0.20 for open weight models and $0.77 for closed source models.
Why do open source models take longer per task?
On Finance Agent v2, the top 5 open weight models take a median of 14.8 minutes per task and the top 5 closed source models take 3.3 minutes. Time per task depends on the number of tool calls, the reasoning effort setting and the speed of the inference provider. Vals AI ran the GLM models at max reasoning effort through Fireworks AI.
Which financial tasks show the largest difference between open and closed source models?
Financial Modeling shows the largest difference on Finance Agent v2 at 9.1 points, followed by Market Analysis at 8.1 points and General Qualitative Analysis at 7.1 points. On the Excel Modeling Benchmark, which tests model building in Excel, the difference is 10.3 points.
Related reading
- Best open source models for financial analysis: ranking, licenses and hardware requirements
- How to access financial statements with ChatGPT or Claude
- How to analyze stocks with AI
- Financial Analysis Prompt Library
- Vals AI Finance Agent v2
- Vals AI Excel Modeling Benchmark
- Vals AI Tax Agent Bench
Last updated: 25 September 2026. Benchmark scores, prices and model versions change with each release. Educational content for general information about AI models.