Best Open Source Models for Financial Analysis in 2026

The open weight models with the highest financial analysis scores in September 2026 are GLM 5.3 Flash (57.9% on Vals AI Finance Agent v2), MiMo V2.6 Pro (57.3), MiMo V2.6 Flash (56.3) and GLM 5.3 (55.8). The top proprietary models score between 58 and 61%. Ling 3.0 Flash Fin from Ant Group is the highest scoring model trained specifically on financial data, and Qwen3.8 27B is the highest scoring model that runs on a single GPU at BF16 or FP8.

Best open source models for financial analysis ranked

The ranking below uses Finance Agent v2 from Vals AI, last updated 23 September 2026. The benchmark gives each model 927 expert reviewed questions based on SEC filings of public companies. Each task has a two hour limit. Each model has six tools: EDGAR search, web search, an HTML page parser, a retrieval tool over fetched pages, a calculator and a price history tool. Scores are the mean of three runs on the private test split of 450 questions. Grading uses partial credit, and an answer scores above zero only when it includes every figure that the answer depends on. Sizes and licenses come from each model's Hugging Face repository.

Rank among open weight modelsModelDeveloperReleasedParameters (total / active)LicenseFinance Agent v2
1GLM 5.3 FlashZ.ai26 Aug 2026320B / 18BMIT57.9%
2MiMo V2.6 ProXiaomi22 Sep 20261.02T / 42BMIT57.3%
3MiMo V2.6 FlashXiaomi22 Sep 2026309B / 15BMIT56.3%
4GLM 5.3Z.ai18 Aug 2026753BGLM 5.3 License55.8%
5Hy4 PreviewTencent28 Aug 2026770B / 49BApache 2.055.1%
6Ling 3.0 Flash FinAnt GroupSep 2026124B / 5.1BMIT54.9%
7Kimi K3Moonshot AI16 Jul 20262.8T / 104BKimi K3 License54.4%
8DeepSeek V4.1 FlashDeepSeek10 Sep 2026552B / 8B to 16BMIT53.5%
9Qwen3.8 MaxAlibaba Qwen3 Aug 20262.4T / 95BQwen3.8 Max License50.6%
10DeepSeek V4 Pro 0813DeepSeek13 Aug 20261.7TMIT50.4%
11GLM 5.2Z.aiJun 2026744B / 40BMIT49.7%
12DeepSeek V4 Flash 0731DeepSeek31 Jul 2026284B / 13BMIT49.5%
13Qwen3.8 27BAlibaba Qwen14 Aug 202627B, denseApache 2.048.6%
14MiniMax M3MiniMax31 May 2026428B / 23BMiniMax Community License48.3%

On the same leaderboard, Gemini 3.8 Flash scores 61.4%, Muse Spark 1.2 scores 60.6% and Claude Opus 5 scores 58.6%. The best proprietary model scores 3.6 percentage points above the best open weight model. Since the 11 September 2026 update, four open weight models entered the top six: MiMo V2.6 Pro, MiMo V2.6 Flash, Hy4 Preview and Ling 3.0 Flash Fin.

GLM 5.3 Flash and GLM 5.3

GLM 5.3 Flash is a mixture of experts model from Z.ai, released under the MIT license. It has 320 billion parameters, and 18 billion of them are active per token. It has the highest score of all 68 models in the Disclosure Analysis category of Finance Agent v2, at 71.5%. This category tracks changes in MD&A language, KPI definitions and segment reporting across several annual filings. The model card describes a reasoning_effort setting with low, high and max levels. It lists deployment through vLLM, SGLang and Transformers. GLM 5.3 is the larger model at 753 billion parameters. It scores 55.8% and is released under a separate GLM 5.3 License.

MiMo V2.6 Pro and MiMo V2.6 Flash

Xiaomi released the weights of MiMo V2.6 Pro and MiMo V2.6 Flash on Hugging Face in September 2026 under the MIT license. MiMo V2.6 Pro has 1.02 trillion parameters with 42 billion active, and MiMo V2.6 Flash has 309 billion parameters with 15 billion active. Both models have a context window of 1,048,576 tokens. MiMo V2.6 Pro scores 57.3% on Finance Agent v2, 0.5 points below GLM 5.3 Flash, and has the highest open weight score in the General Quantitative Analysis category at 79.8%. MiMo V2.6 Flash scores 56.3% and has the highest open weight score in the Comparables category at 49.0%. On the Vals AI run, MiMo V2.6 Pro costs $0.20 per task and MiMo V2.6 Flash costs $0.07 per task through Xiaomi's API.

Hy4 Preview

Hy4 Preview from Tencent has 770 billion parameters with 49 billion active and a context window of 1 million tokens. Tencent released it on 28 August 2026 under the Apache 2.0 license. It scores 55.1% on Finance Agent v2 and 63.7% on Tax Agent Bench. Its time per task on Finance Agent v2 is 23.7 minutes, the longest of the top ten open weight models.

Ling 3.0 Flash Fin

Ling 3.0 Flash Fin from Ant Group is Ling 3.0 Flash with continued training on financial data, developed with financial institutions. It has 124 billion parameters with 5.1 billion active, a context window of 256,000 tokens and the MIT license. It scores 54.9% on Finance Agent v2 at $0.04 per task, the lowest cost per task of the 20 highest scoring models on the leaderboard. The general purpose Ling 3.0 Flash scores 30.3% on the same benchmark, 24.6 points lower. Ant Group released FinFIRST alongside the model, a benchmark of 123 expert authored investment research tasks.

Kimi K3

Kimi K3 from Moonshot AI has 2.8 trillion parameters with 104 billion active. It selects 16 of 896 experts per token and has a context window of 1,048,576 tokens. The model card reports three scores. It scores 54.4 on Finance Agent v2. It scores 71.6 on CorpFin v2, a benchmark on credit agreements that typically exceed 200 pages. It scores 94.5 on MCPMark Verified, a test of tool use over MCP servers. The weights are available on Hugging Face in Safetensors format.

DeepSeek V4.1 Flash and DeepSeek V4 Pro

DeepSeek released V4.1 Flash on 10 September 2026 under the MIT license. The model has 552 billion parameters in total. It activates 8 billion parameters during prefill and 16 billion during decoding. It reads up to 1 million tokens and accepts images as input. Its reasoning effort is an integer from 1 to 100, and this setting controls the balance of cost and accuracy per request. The earlier DeepSeek V4 scored 60.39% on Finance Agent v1.1 in June 2026. That score placed it fourth among all models on that version of the benchmark. DeepSeek V4 Pro 0813 is the 1.7 trillion parameter release that replaces the V4 Pro preview.

Qwen3.8 Max and Qwen3.8 27B

Qwen3.8 Max has 2.4 trillion parameters with 95 billion active. It runs in thinking mode for every request and accepts text input. Qwen3.8 27B is a dense model released under Apache 2.0. It has a native context of 262,144 tokens, extensible to 1 million. At 27 billion parameters, it is the only model in this ranking that can be served from one 80 GB GPU at BF16 or FP8. Its score of 48.6% is 1.8 points below DeepSeek V4 Pro 0813. DeepSeek V4 Pro 0813 has 63 times the parameter count of Qwen3.8 27B.

MiniMax M3

MiniMax M3 has about 428 billion parameters with about 23 billion active and a 1 million token context. It uses MiniMax Sparse Attention. The model card reports that this makes it 9 times faster at prefill and 15 times faster at decoding than MiniMax M2 at full context. It ranks 40th of 68 on Finance Agent v2 at 48.3%. Vals AI lists hosted access prices of $0.60 per million input tokens and $2.40 per million output tokens.

Open source vs open weight: license terms for commercial use

All models in the ranking publish their weights. A model that publishes its weights is an open weight model; an open source model also publishes its training data and code. The license sets the terms under which a bank, fund or data vendor can deploy a model, and whether a separate agreement with the developer applies.

LicenseModelsCommercial useConditions
MITGLM 5.3 Flash, GLM 5.2, MiMo V2.6 Pro, MiMo V2.6 Flash, Ling 3.0 Flash Fin, DeepSeek V4.1 Flash, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731PermittedKeep the copyright and permission notice
Apache 2.0Hy4 Preview, Qwen3.8 27B, gpt oss 120b, gpt oss 20bPermittedKeep the license and notices, state changes made to the files
GLM 5.3 LicenseGLM 5.3PermittedA company offering the model as a service with more than $10 billion in revenue over 12 months passes a security review by Z.ai first
Kimi K3 LicenseKimi K3PermittedA separate agreement for model as a service businesses above $20 million in revenue over 12 months; "Kimi K3" displayed in products above 100 million monthly active users or $20 million monthly revenue; internal use is exempt
Qwen3.8 Max LicenseQwen3.8 MaxPermittedPermission from Qwen for model as a service or AI work assistant businesses above $50 million in revenue over 12 months; model name displayed above 100 million monthly active users or $20 million monthly revenue
MiniMax Community LicenseMiniMax M3Permitted with noticeOne time notice to MiniMax below $20 million annual revenue, prior written authorization above it, "Built with MiniMax M3" displayed, listed prohibited uses

Internal research means the model runs on the firm's own infrastructure and serves only its own staff. For this use, the Kimi K3 License exempts the firm from its revenue and display clauses. The Qwen3.8 Max License permits internal use of the model as a service. A data vendor that gives clients access to the model through an API falls under the model as a service definitions in the GLM 5.3, Kimi K3 and Qwen3.8 Max licenses. Five of the six highest scoring open weight models, all except GLM 5.3, use MIT or Apache 2.0.

What benchmarks exist to test AI models for financial analysis

Two types of benchmarks test AI models on finance. Agent benchmarks give the model tools, such as EDGAR search or a spreadsheet, and score the finished answer or file. Q&A benchmarks give the model a question and the source text or table, and score the answer.

Agent benchmarks for finance

BenchmarkMaintainerWhat it testsSize
Finance Agent v2Vals AIAnalyst tasks on SEC filings with EDGAR search, web search, a calculator and price history927 questions
Finance Agent v1.1Vals AIEarlier version of Finance Agent with retrieval, beat or miss and trend questions537 questions
Excel Modeling BenchmarkVals AILBO, DCF, M&A, operating and comparable company models built in Excel103 tasks
Tax Agent BenchVals AIResearch questions on US corporate tax, answered with multi step research391 questions
CorpFin v2Vals AITerm extraction and numeric reasoning over credit agreements of 200 pages or more858 questions
TaxEval v2Vals AITax questions with reference answers written by Vals AISee benchmark page
MortgageTaxVals AIReading tax certificates provided as imagesSee benchmark page

Q&A benchmarks for finance

BenchmarkMaintainerWhat it testsSize
FinanceBenchPatronus AIQuestions about public companies answered from their SEC filings10,231 questions
FinQAResearch paperMulti step calculations over tables and text from earnings reports8,281 questions
ConvFinQAResearch paperCalculations over earnings reports across several turns of a conversation3,892 conversations
TAT-QAResearch paperArithmetic and comparisons over tables combined with text from financial reports16,552 questions
DocFinQAResearch paperFinQA questions with the full report as context, about 123,000 words per document7,437 questions
BizBenchResearch paperQuantitative reasoning on business and finance, including code generation8 tasks
FinBenThe FinAIInformation extraction, text analysis, question answering, forecasting, risk management and stock trading36 datasets, 24 tasks

Agent benchmarks measure the full workflow: finding the filing, reading it and calculating the answer. Q&A benchmarks measure the reading and calculation steps on a given document. The ranking in this article uses Finance Agent v2 because its tasks use SEC filings and tools, which matches how an analyst uses a model connected to filing data.

How the Finance Agent benchmark tests financial analysis

Finance Agent v2 sorts its questions into nine categories. The categories are based on the work of a second or third year investment banking analyst. The leader score is the highest score any model reaches in each category.

CategoryExample task from the benchmarkLeader score
General Qualitative AnalysisCompare Walmart, Costco and Target capital allocation across capex, dividends, buybacks and debt83.4%
General Quantitative AnalysisCalculate the difference in days inventory outstanding between Home Depot and Lowe's for FY202481.8%
Earnings AnalysisCompare Rapid7's Q3 2025 actuals against prior revenue, non GAAP operating income and ARR guidance81.4%
Market AnalysisMeasure a stock's reaction to an announced divestiture and relate it to the stated use of proceeds78.5%
Disclosure AnalysisTrack Boeing's segment reporting and 787 cost recovery disclosures across FY2022 to FY2024 10-K filings71.5% (GLM 5.3 Flash)
AdjustmentsReconcile Honeywell's GAAP operating income to segment profit across annual releases56.3%
ComparablesRank major US banks by excess CET1 ratio over their regulatory minimums52.0%
PrecedentsExtract EV/EBITDA multiples for industrial distribution acquisitions from S-4 filings36.4%
Financial ModelingBuild a DCF, LBO or accretion and dilution model from historical ratios in the filings34.5%

For retrieval and summarization of a filing, the category leaders score above 70%. Tasks that reconcile figures across several documents, or build a model from those figures, score between 34 and 56%. The All Pass metric requires every check in an answer to pass. Under this metric, the highest score on the leaderboard is 50.88%. Vals AI also reports on tool use. Higher scoring models make more tool calls, and most of those calls go to the calculator after the model locates the source figures. Lower scoring models make more exploratory web search and retrieval calls.

Based on these results, a finance team can use the model to retrieve and structure filing data. An analyst then reviews any multi step valuation output against the source filings.

How these models score against the closed source models on the same benchmarks is covered in Open source vs closed source AI models for financial analysis.

Hardware needed to run open source models locally

Memory for the weights equals the parameter count times the bytes per parameter. This is 2 bytes at BF16, 1 byte at FP8 and about half a byte at 4 bit quantization. The key value cache for long filings and concurrent users needs additional memory.

ModelWeights at FP8Typical deployment
Qwen3.8 27Babout 27 GBOne 80 GB GPU at BF16 or FP8
gpt oss 120babout 61 GB as released in MXFP4One 80 GB GPU
Ling 3.0 Flash Finabout 124 GBTwo 80 GB GPUs or one GPU of 141 GB or more
DeepSeek V4 Flash 0731about 284 GBOne node with 8 GPUs of 80 GB or more
MiMo V2.6 Flashabout 309 GBOne node with 8 GPUs of 80 GB or more
GLM 5.3 Flashabout 320 GBOne node with 8 GPUs of 80 GB or more
MiniMax M3about 428 GBOne node with 8 GPUs of 80 GB or more
GLM 5.3, DeepSeek V4.1 Flash, Hy4 Preview550 to 770 GBOne node with 8 GPUs of 141 GB or more
MiMo V2.6 Proabout 1.02 TBOne node with 8 GPUs of 192 GB, or several nodes
DeepSeek V4 Pro 0813, Qwen3.8 Max, Kimi K31.7 to 2.8 TBSeveral nodes

Mixture of experts models keep every expert in memory and compute only the active parameters for each token. GLM 5.3 Flash has 18 billion active parameters. It generates tokens at a speed closer to an 18 billion parameter dense model and needs the memory of a 320 billion parameter model. Ling 3.0 Flash Fin has 5.1 billion active parameters, the fewest of the top ten open weight models. Firms can also keep the weights under their control on dedicated instances at a cloud provider, inside their own cloud account.

FAQ

What is the best open source model for financial analysis?

GLM 5.3 Flash from Z.ai has the highest open weight score on Vals AI Finance Agent v2 as of 23 September 2026, at 57.9%. The top proprietary model scores 61.4%. GLM 5.3 Flash is released under the MIT license and has 320 billion parameters with 18 billion active. It has the highest score of all 68 models in the Disclosure Analysis category. MiMo V2.6 Pro, MiMo V2.6 Flash and GLM 5.3 score 57.3, 56.3 and 55.8%.

Is there an open source model trained for finance?

Ling 3.0 Flash Fin from Ant Group is trained on financial data on top of the general purpose Ling 3.0 Flash. It is released under the MIT license with 124 billion parameters and 5.1 billion active. It scores 54.9% on Finance Agent v2, 24.6 points above the general purpose Ling 3.0 Flash, and costs $0.04 per task on the Vals AI run.

Can an open source model run on a single GPU for financial analysis?

Qwen3.8 27B is a dense 27 billion parameter model under Apache 2.0. It fits on one 80 GB GPU and scores 48.6% on Finance Agent v2. This places it 13th among open weight models and about 2 points below models with over a trillion parameters. gpt oss 120b from OpenAI also runs on one 80 GB GPU in its MXFP4 release. Ling 3.0 Flash Fin at 124 billion parameters fits on one GPU of 141 GB or more at FP8. Larger mixture of experts models need a multi GPU node.

Are open source models good enough for financial analysis compared with GPT, Claude and Gemini?

On Finance Agent v2, the best open weight model scores 3.6 percentage points below the best proprietary model. Several open weight models score above proprietary models from the first half of 2026, including Claude Opus 4.7 at 51.5%. Across all models, the category leaders score above 70% on tasks that retrieve and summarize filings. The category leaders score below 37% on financial modeling and precedent transaction tasks.

Can banks and funds use these models commercially?

Models under MIT and Apache 2.0 allow commercial use when the user keeps the license notice. These include GLM 5.3 Flash, MiMo V2.6 Pro and Flash, Hy4 Preview, Ling 3.0 Flash Fin, the DeepSeek V4 family and Qwen3.8 27B. Kimi K3, Qwen3.8 Max, GLM 5.3 and MiniMax M3 use custom licenses with revenue thresholds, attribution clauses or notice requirements. Most of these terms apply to firms that offer the model to third parties as a service. The license file in each Hugging Face repository contains the full legal terms.

Why run an open weight model in place of a hosted API?

Self hosting keeps prompts, documents and outputs inside the firm's own infrastructure. This supports data residency and confidentiality requirements for material non public information, client portfolios and draft deal documents. Self hosting also fixes the model version, so a validated workflow produces the same behaviour until the firm upgrades the model. It also allows fine tuning on internal research.

How does an open source model get access to SEC filings?

The model accesses filings through tools. The SEC-API.io MCP server provides EDGAR filings, XBRL financial statements, full text search, insider trades and institutional holdings as tools. Any MCP compatible agent framework can attach these tools to a self hosted model. For batch pipelines, the SEC-API.io REST APIs return the same data as JSON or extracted text.

Connecting a model to SEC data

Benchmarks and model sources

Last updated: 25 September 2026. Benchmark scores and license terms change with each model release. Check the Vals AI leaderboard and the license file in each Hugging Face repository before deployment. Educational content for general information about AI models and their licenses.