Best Open Source Models for Financial Analysis in 2026
The open weight models with the highest financial analysis scores in September 2026 are GLM 5.3 Flash (57.9% on Vals AI Finance Agent v2), MiMo V2.6 Pro (57.3), MiMo V2.6 Flash (56.3) and GLM 5.3 (55.8). The top proprietary models score between 58 and 61%. Ling 3.0 Flash Fin from Ant Group is the highest scoring model trained specifically on financial data, and Qwen3.8 27B is the highest scoring model that runs on a single GPU at BF16 or FP8.
On this page:
- Best open source models for financial analysis ranked
- GLM 5.3 Flash and GLM 5.3
- MiMo V2.6 Pro and MiMo V2.6 Flash
- Hy4 Preview
- Ling 3.0 Flash Fin
- Kimi K3
- DeepSeek V4.1 Flash and DeepSeek V4 Pro
- Qwen3.8 Max and Qwen3.8 27B
- MiniMax M3
- Open source vs open weight: license terms for commercial use
- What benchmarks exist to test AI models for financial analysis
- Agent benchmarks for finance
- Q&A benchmarks for finance
- How the Finance Agent benchmark tests financial analysis
- Hardware needed to run open source models locally
- FAQ
- Related reading
Best open source models for financial analysis ranked
The ranking below uses Finance Agent v2 from Vals AI, last updated 23 September 2026. The benchmark gives each model 927 expert reviewed questions based on SEC filings of public companies. Each task has a two hour limit. Each model has six tools: EDGAR search, web search, an HTML page parser, a retrieval tool over fetched pages, a calculator and a price history tool. Scores are the mean of three runs on the private test split of 450 questions. Grading uses partial credit, and an answer scores above zero only when it includes every figure that the answer depends on. Sizes and licenses come from each model's Hugging Face repository.
| Rank among open weight models | Model | Developer | Released | Parameters (total / active) | License | Finance Agent v2 |
|---|---|---|---|---|---|---|
| 1 | GLM 5.3 Flash | Z.ai | 26 Aug 2026 | 320B / 18B | MIT | 57.9% |
| 2 | MiMo V2.6 Pro | Xiaomi | 22 Sep 2026 | 1.02T / 42B | MIT | 57.3% |
| 3 | MiMo V2.6 Flash | Xiaomi | 22 Sep 2026 | 309B / 15B | MIT | 56.3% |
| 4 | GLM 5.3 | Z.ai | 18 Aug 2026 | 753B | GLM 5.3 License | 55.8% |
| 5 | Hy4 Preview | Tencent | 28 Aug 2026 | 770B / 49B | Apache 2.0 | 55.1% |
| 6 | Ling 3.0 Flash Fin | Ant Group | Sep 2026 | 124B / 5.1B | MIT | 54.9% |
| 7 | Kimi K3 | Moonshot AI | 16 Jul 2026 | 2.8T / 104B | Kimi K3 License | 54.4% |
| 8 | DeepSeek V4.1 Flash | DeepSeek | 10 Sep 2026 | 552B / 8B to 16B | MIT | 53.5% |
| 9 | Qwen3.8 Max | Alibaba Qwen | 3 Aug 2026 | 2.4T / 95B | Qwen3.8 Max License | 50.6% |
| 10 | DeepSeek V4 Pro 0813 | DeepSeek | 13 Aug 2026 | 1.7T | MIT | 50.4% |
| 11 | GLM 5.2 | Z.ai | Jun 2026 | 744B / 40B | MIT | 49.7% |
| 12 | DeepSeek V4 Flash 0731 | DeepSeek | 31 Jul 2026 | 284B / 13B | MIT | 49.5% |
| 13 | Qwen3.8 27B | Alibaba Qwen | 14 Aug 2026 | 27B, dense | Apache 2.0 | 48.6% |
| 14 | MiniMax M3 | MiniMax | 31 May 2026 | 428B / 23B | MiniMax Community License | 48.3% |
On the same leaderboard, Gemini 3.8 Flash scores 61.4%, Muse Spark 1.2 scores 60.6% and Claude Opus 5 scores 58.6%. The best proprietary model scores 3.6 percentage points above the best open weight model. Since the 11 September 2026 update, four open weight models entered the top six: MiMo V2.6 Pro, MiMo V2.6 Flash, Hy4 Preview and Ling 3.0 Flash Fin.
GLM 5.3 Flash and GLM 5.3
GLM 5.3 Flash is a mixture of experts model from Z.ai, released under the MIT license. It has 320 billion parameters, and 18 billion of them are active per token. It has the highest score of all 68 models in the Disclosure Analysis category of Finance Agent v2, at 71.5%. This category tracks changes in MD&A language, KPI definitions and segment reporting across several annual filings. The model card describes a reasoning_effort setting with low, high and max levels. It lists deployment through vLLM, SGLang and Transformers. GLM 5.3 is the larger model at 753 billion parameters. It scores 55.8% and is released under a separate GLM 5.3 License.
MiMo V2.6 Pro and MiMo V2.6 Flash
Xiaomi released the weights of MiMo V2.6 Pro and MiMo V2.6 Flash on Hugging Face in September 2026 under the MIT license. MiMo V2.6 Pro has 1.02 trillion parameters with 42 billion active, and MiMo V2.6 Flash has 309 billion parameters with 15 billion active. Both models have a context window of 1,048,576 tokens. MiMo V2.6 Pro scores 57.3% on Finance Agent v2, 0.5 points below GLM 5.3 Flash, and has the highest open weight score in the General Quantitative Analysis category at 79.8%. MiMo V2.6 Flash scores 56.3% and has the highest open weight score in the Comparables category at 49.0%. On the Vals AI run, MiMo V2.6 Pro costs $0.20 per task and MiMo V2.6 Flash costs $0.07 per task through Xiaomi's API.
Hy4 Preview
Hy4 Preview from Tencent has 770 billion parameters with 49 billion active and a context window of 1 million tokens. Tencent released it on 28 August 2026 under the Apache 2.0 license. It scores 55.1% on Finance Agent v2 and 63.7% on Tax Agent Bench. Its time per task on Finance Agent v2 is 23.7 minutes, the longest of the top ten open weight models.
Ling 3.0 Flash Fin
Ling 3.0 Flash Fin from Ant Group is Ling 3.0 Flash with continued training on financial data, developed with financial institutions. It has 124 billion parameters with 5.1 billion active, a context window of 256,000 tokens and the MIT license. It scores 54.9% on Finance Agent v2 at $0.04 per task, the lowest cost per task of the 20 highest scoring models on the leaderboard. The general purpose Ling 3.0 Flash scores 30.3% on the same benchmark, 24.6 points lower. Ant Group released FinFIRST alongside the model, a benchmark of 123 expert authored investment research tasks.
Kimi K3
Kimi K3 from Moonshot AI has 2.8 trillion parameters with 104 billion active. It selects 16 of 896 experts per token and has a context window of 1,048,576 tokens. The model card reports three scores. It scores 54.4 on Finance Agent v2. It scores 71.6 on CorpFin v2, a benchmark on credit agreements that typically exceed 200 pages. It scores 94.5 on MCPMark Verified, a test of tool use over MCP servers. The weights are available on Hugging Face in Safetensors format.
DeepSeek V4.1 Flash and DeepSeek V4 Pro
DeepSeek released V4.1 Flash on 10 September 2026 under the MIT license. The model has 552 billion parameters in total. It activates 8 billion parameters during prefill and 16 billion during decoding. It reads up to 1 million tokens and accepts images as input. Its reasoning effort is an integer from 1 to 100, and this setting controls the balance of cost and accuracy per request. The earlier DeepSeek V4 scored 60.39% on Finance Agent v1.1 in June 2026. That score placed it fourth among all models on that version of the benchmark. DeepSeek V4 Pro 0813 is the 1.7 trillion parameter release that replaces the V4 Pro preview.
Qwen3.8 Max and Qwen3.8 27B
Qwen3.8 Max has 2.4 trillion parameters with 95 billion active. It runs in thinking mode for every request and accepts text input. Qwen3.8 27B is a dense model released under Apache 2.0. It has a native context of 262,144 tokens, extensible to 1 million. At 27 billion parameters, it is the only model in this ranking that can be served from one 80 GB GPU at BF16 or FP8. Its score of 48.6% is 1.8 points below DeepSeek V4 Pro 0813. DeepSeek V4 Pro 0813 has 63 times the parameter count of Qwen3.8 27B.
MiniMax M3
MiniMax M3 has about 428 billion parameters with about 23 billion active and a 1 million token context. It uses MiniMax Sparse Attention. The model card reports that this makes it 9 times faster at prefill and 15 times faster at decoding than MiniMax M2 at full context. It ranks 40th of 68 on Finance Agent v2 at 48.3%. Vals AI lists hosted access prices of $0.60 per million input tokens and $2.40 per million output tokens.
Open source vs open weight: license terms for commercial use
All models in the ranking publish their weights. A model that publishes its weights is an open weight model; an open source model also publishes its training data and code. The license sets the terms under which a bank, fund or data vendor can deploy a model, and whether a separate agreement with the developer applies.
| License | Models | Commercial use | Conditions |
|---|---|---|---|
| MIT | GLM 5.3 Flash, GLM 5.2, MiMo V2.6 Pro, MiMo V2.6 Flash, Ling 3.0 Flash Fin, DeepSeek V4.1 Flash, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731 | Permitted | Keep the copyright and permission notice |
| Apache 2.0 | Hy4 Preview, Qwen3.8 27B, gpt oss 120b, gpt oss 20b | Permitted | Keep the license and notices, state changes made to the files |
| GLM 5.3 License | GLM 5.3 | Permitted | A company offering the model as a service with more than $10 billion in revenue over 12 months passes a security review by Z.ai first |
| Kimi K3 License | Kimi K3 | Permitted | A separate agreement for model as a service businesses above $20 million in revenue over 12 months; "Kimi K3" displayed in products above 100 million monthly active users or $20 million monthly revenue; internal use is exempt |
| Qwen3.8 Max License | Qwen3.8 Max | Permitted | Permission from Qwen for model as a service or AI work assistant businesses above $50 million in revenue over 12 months; model name displayed above 100 million monthly active users or $20 million monthly revenue |
| MiniMax Community License | MiniMax M3 | Permitted with notice | One time notice to MiniMax below $20 million annual revenue, prior written authorization above it, "Built with MiniMax M3" displayed, listed prohibited uses |
Internal research means the model runs on the firm's own infrastructure and serves only its own staff. For this use, the Kimi K3 License exempts the firm from its revenue and display clauses. The Qwen3.8 Max License permits internal use of the model as a service. A data vendor that gives clients access to the model through an API falls under the model as a service definitions in the GLM 5.3, Kimi K3 and Qwen3.8 Max licenses. Five of the six highest scoring open weight models, all except GLM 5.3, use MIT or Apache 2.0.
What benchmarks exist to test AI models for financial analysis
Two types of benchmarks test AI models on finance. Agent benchmarks give the model tools, such as EDGAR search or a spreadsheet, and score the finished answer or file. Q&A benchmarks give the model a question and the source text or table, and score the answer.
Agent benchmarks for finance
| Benchmark | Maintainer | What it tests | Size |
|---|---|---|---|
| Finance Agent v2 | Vals AI | Analyst tasks on SEC filings with EDGAR search, web search, a calculator and price history | 927 questions |
| Finance Agent v1.1 | Vals AI | Earlier version of Finance Agent with retrieval, beat or miss and trend questions | 537 questions |
| Excel Modeling Benchmark | Vals AI | LBO, DCF, M&A, operating and comparable company models built in Excel | 103 tasks |
| Tax Agent Bench | Vals AI | Research questions on US corporate tax, answered with multi step research | 391 questions |
| CorpFin v2 | Vals AI | Term extraction and numeric reasoning over credit agreements of 200 pages or more | 858 questions |
| TaxEval v2 | Vals AI | Tax questions with reference answers written by Vals AI | See benchmark page |
| MortgageTax | Vals AI | Reading tax certificates provided as images | See benchmark page |
Q&A benchmarks for finance
| Benchmark | Maintainer | What it tests | Size |
|---|---|---|---|
| FinanceBench | Patronus AI | Questions about public companies answered from their SEC filings | 10,231 questions |
| FinQA | Research paper | Multi step calculations over tables and text from earnings reports | 8,281 questions |
| ConvFinQA | Research paper | Calculations over earnings reports across several turns of a conversation | 3,892 conversations |
| TAT-QA | Research paper | Arithmetic and comparisons over tables combined with text from financial reports | 16,552 questions |
| DocFinQA | Research paper | FinQA questions with the full report as context, about 123,000 words per document | 7,437 questions |
| BizBench | Research paper | Quantitative reasoning on business and finance, including code generation | 8 tasks |
| FinBen | The FinAI | Information extraction, text analysis, question answering, forecasting, risk management and stock trading | 36 datasets, 24 tasks |
Agent benchmarks measure the full workflow: finding the filing, reading it and calculating the answer. Q&A benchmarks measure the reading and calculation steps on a given document. The ranking in this article uses Finance Agent v2 because its tasks use SEC filings and tools, which matches how an analyst uses a model connected to filing data.
How the Finance Agent benchmark tests financial analysis
Finance Agent v2 sorts its questions into nine categories. The categories are based on the work of a second or third year investment banking analyst. The leader score is the highest score any model reaches in each category.
| Category | Example task from the benchmark | Leader score |
|---|---|---|
| General Qualitative Analysis | Compare Walmart, Costco and Target capital allocation across capex, dividends, buybacks and debt | 83.4% |
| General Quantitative Analysis | Calculate the difference in days inventory outstanding between Home Depot and Lowe's for FY2024 | 81.8% |
| Earnings Analysis | Compare Rapid7's Q3 2025 actuals against prior revenue, non GAAP operating income and ARR guidance | 81.4% |
| Market Analysis | Measure a stock's reaction to an announced divestiture and relate it to the stated use of proceeds | 78.5% |
| Disclosure Analysis | Track Boeing's segment reporting and 787 cost recovery disclosures across FY2022 to FY2024 10-K filings | 71.5% (GLM 5.3 Flash) |
| Adjustments | Reconcile Honeywell's GAAP operating income to segment profit across annual releases | 56.3% |
| Comparables | Rank major US banks by excess CET1 ratio over their regulatory minimums | 52.0% |
| Precedents | Extract EV/EBITDA multiples for industrial distribution acquisitions from S-4 filings | 36.4% |
| Financial Modeling | Build a DCF, LBO or accretion and dilution model from historical ratios in the filings | 34.5% |
For retrieval and summarization of a filing, the category leaders score above 70%. Tasks that reconcile figures across several documents, or build a model from those figures, score between 34 and 56%. The All Pass metric requires every check in an answer to pass. Under this metric, the highest score on the leaderboard is 50.88%. Vals AI also reports on tool use. Higher scoring models make more tool calls, and most of those calls go to the calculator after the model locates the source figures. Lower scoring models make more exploratory web search and retrieval calls.
Based on these results, a finance team can use the model to retrieve and structure filing data. An analyst then reviews any multi step valuation output against the source filings.
How these models score against the closed source models on the same benchmarks is covered in Open source vs closed source AI models for financial analysis.
Hardware needed to run open source models locally
Memory for the weights equals the parameter count times the bytes per parameter. This is 2 bytes at BF16, 1 byte at FP8 and about half a byte at 4 bit quantization. The key value cache for long filings and concurrent users needs additional memory.
| Model | Weights at FP8 | Typical deployment |
|---|---|---|
| Qwen3.8 27B | about 27 GB | One 80 GB GPU at BF16 or FP8 |
| gpt oss 120b | about 61 GB as released in MXFP4 | One 80 GB GPU |
| Ling 3.0 Flash Fin | about 124 GB | Two 80 GB GPUs or one GPU of 141 GB or more |
| DeepSeek V4 Flash 0731 | about 284 GB | One node with 8 GPUs of 80 GB or more |
| MiMo V2.6 Flash | about 309 GB | One node with 8 GPUs of 80 GB or more |
| GLM 5.3 Flash | about 320 GB | One node with 8 GPUs of 80 GB or more |
| MiniMax M3 | about 428 GB | One node with 8 GPUs of 80 GB or more |
| GLM 5.3, DeepSeek V4.1 Flash, Hy4 Preview | 550 to 770 GB | One node with 8 GPUs of 141 GB or more |
| MiMo V2.6 Pro | about 1.02 TB | One node with 8 GPUs of 192 GB, or several nodes |
| DeepSeek V4 Pro 0813, Qwen3.8 Max, Kimi K3 | 1.7 to 2.8 TB | Several nodes |
Mixture of experts models keep every expert in memory and compute only the active parameters for each token. GLM 5.3 Flash has 18 billion active parameters. It generates tokens at a speed closer to an 18 billion parameter dense model and needs the memory of a 320 billion parameter model. Ling 3.0 Flash Fin has 5.1 billion active parameters, the fewest of the top ten open weight models. Firms can also keep the weights under their control on dedicated instances at a cloud provider, inside their own cloud account.
FAQ
What is the best open source model for financial analysis?
GLM 5.3 Flash from Z.ai has the highest open weight score on Vals AI Finance Agent v2 as of 23 September 2026, at 57.9%. The top proprietary model scores 61.4%. GLM 5.3 Flash is released under the MIT license and has 320 billion parameters with 18 billion active. It has the highest score of all 68 models in the Disclosure Analysis category. MiMo V2.6 Pro, MiMo V2.6 Flash and GLM 5.3 score 57.3, 56.3 and 55.8%.
Is there an open source model trained for finance?
Ling 3.0 Flash Fin from Ant Group is trained on financial data on top of the general purpose Ling 3.0 Flash. It is released under the MIT license with 124 billion parameters and 5.1 billion active. It scores 54.9% on Finance Agent v2, 24.6 points above the general purpose Ling 3.0 Flash, and costs $0.04 per task on the Vals AI run.
Can an open source model run on a single GPU for financial analysis?
Qwen3.8 27B is a dense 27 billion parameter model under Apache 2.0. It fits on one 80 GB GPU and scores 48.6% on Finance Agent v2. This places it 13th among open weight models and about 2 points below models with over a trillion parameters. gpt oss 120b from OpenAI also runs on one 80 GB GPU in its MXFP4 release. Ling 3.0 Flash Fin at 124 billion parameters fits on one GPU of 141 GB or more at FP8. Larger mixture of experts models need a multi GPU node.
Are open source models good enough for financial analysis compared with GPT, Claude and Gemini?
On Finance Agent v2, the best open weight model scores 3.6 percentage points below the best proprietary model. Several open weight models score above proprietary models from the first half of 2026, including Claude Opus 4.7 at 51.5%. Across all models, the category leaders score above 70% on tasks that retrieve and summarize filings. The category leaders score below 37% on financial modeling and precedent transaction tasks.
Can banks and funds use these models commercially?
Models under MIT and Apache 2.0 allow commercial use when the user keeps the license notice. These include GLM 5.3 Flash, MiMo V2.6 Pro and Flash, Hy4 Preview, Ling 3.0 Flash Fin, the DeepSeek V4 family and Qwen3.8 27B. Kimi K3, Qwen3.8 Max, GLM 5.3 and MiniMax M3 use custom licenses with revenue thresholds, attribution clauses or notice requirements. Most of these terms apply to firms that offer the model to third parties as a service. The license file in each Hugging Face repository contains the full legal terms.
Why run an open weight model in place of a hosted API?
Self hosting keeps prompts, documents and outputs inside the firm's own infrastructure. This supports data residency and confidentiality requirements for material non public information, client portfolios and draft deal documents. Self hosting also fixes the model version, so a validated workflow produces the same behaviour until the firm upgrades the model. It also allows fine tuning on internal research.
How does an open source model get access to SEC filings?
The model accesses filings through tools. The SEC-API.io MCP server provides EDGAR filings, XBRL financial statements, full text search, insider trades and institutional holdings as tools. Any MCP compatible agent framework can attach these tools to a self hosted model. For batch pipelines, the SEC-API.io REST APIs return the same data as JSON or extracted text.
Related reading
- Open source vs closed source AI models for financial analysis: where the gap stands across the benchmarks
Connecting a model to SEC data
- MCP server documentation: endpoint and full tool list
- How to access financial statements with ChatGPT or Claude
- How to access SEC financial data in Claude
- How to access SEC financial data in ChatGPT
- Financial Analysis Prompt Library
- API documentation
- Get a free API key
Benchmarks and model sources
- Vals AI Finance Agent v2
- Finance Agent Benchmark paper
- GLM 5.3 Flash on Hugging Face
- MiMo V2.6 models on Hugging Face
- Hy4 Preview on Hugging Face
- Ling 3.0 Flash Fin on Hugging Face
- Kimi K3 on Hugging Face
- DeepSeek V4.1 Flash on Hugging Face
- Qwen3.8 27B on Hugging Face
- MiniMax M3 on Hugging Face
Last updated: 25 September 2026. Benchmark scores and license terms change with each model release. Check the Vals AI leaderboard and the license file in each Hugging Face repository before deployment. Educational content for general information about AI models and their licenses.