The best open-weight time-series foundation models in October 2026 are TimesFM 3.0 from Google, TiRex-2 from NXAI, Toto 2.0 from Datadog, and Chronos-2 from Amazon. TimesFM 3.0 has the best GIFT-Eval scores of the group, but its weights carry a non-commercial licence. For commercial work the choice is between TiRex-2, Toto 2.0 and Chronos-2, all Apache 2.0 and within two percent of each other on point accuracy. Chronos-2 is by far the most used: Hugging Face recorded about 22.7 million downloads in the 30 days to October 1, 2026.
A time-series foundation model is pretrained on many series and forecasts a new one without training on it. Lower scores are better in the table below. A score of 1.000 equals the Seasonal Naive baseline.
Open-Weight Time-Series Model Comparison
| Model | Developer | Parameters | License | Released | GIFT-Eval MASE and CRPS (leaderboard result files) | Hardware to run |
|---|---|---|---|---|---|---|
| TimesFM 3.0 | Google Research | 331M | TimesFM Non-Commercial License v1.0 | August 2026 | 0.667 and 0.456 | Not stated on card |
| TiRex-2 | NXAI | 38.4M active, plus 44.1M in multivariate mode | Apache 2.0 | July 2026 | 0.697 and 0.478 for the zero-shot checkpoint | CPU or CUDA GPU |
| Granite PatchTST-FM r2 | IBM and Rensselaer Polytechnic Institute | 385M | OpenMDW 1.0 | August 2026 | 0.685 and 0.467 | Not stated on card |
| Toto 2.0 2.5B | Datadog | 2.5B (also 4M, 22M, 313M, 1B) | Apache 2.0 | May 2026 | 0.696 and 0.476 | 9.1 GB weights, about 36 ms per forecast on an A100 |
| Chronos-2 | Amazon | 120M | Apache 2.0 | October 2025 | 0.698 and 0.485 | Over 300 series per second on one A10G. CPU supported |
| TimesFM 2.5 | Google Research | 200M | Apache 2.0 | September 2025 | 0.705 and 0.490 | Not stated on card |
| Toto 2.0 22m | Datadog | 22M | Apache 2.0 | May 2026 | 0.719 and 0.496 | 84 MB weights, about 5 ms on an A100 |
| TiRex | NXAI | 35M | NXAI Community License | May 2025 | 0.716 and 0.488 | CUDA GPU, CPU fallback |
| Moirai 2.0 small | Salesforce | 11.4M | CC BY-NC 4.0 | August 2025 | 0.728 and 0.516 | Not stated on card |
| Sundial base | Tsinghua University | 128M | Apache 2.0 | May 2025 | 0.750 and 0.559 | Not stated on card |
The older generation has been left behind. Lag-Llama scores 1.228 on MASE in the same files, worse than Seasonal Naive, and the leaderboard flags it for test data leakage.
GIFT-Eval is run by Salesforce AI Research and covers 97 task configurations. The scores above are aggregated from the leaderboard's published per-task result files, normalised against Seasonal Naive. Amazon's Chronos-2 paper reports a second benchmark, fev-bench, where Chronos-2 had a skill score of 47.3 percent against 42.6 for TiRex and 42.3 for TimesFM 2.5. That paper is from October 2025 and predates the 2026 releases.
Which to Pick for Which Job
| Job | Pick | Reason |
|---|---|---|
| Commercial forecasting with covariates | Chronos-2 or TiRex-2 | Both take past and known future covariates and are Apache 2.0 |
| Research where licence does not matter | TimesFM 3.0 | Best MASE and CRPS in this table |
| Observability and infrastructure metrics | Toto 2.0 | Trained by Datadog on observability data, and tops Datadog's BOOM benchmark |
| CPU or edge devices | Toto 2.0 4m or TiRex-2 | Toto 4m is a 16 MB file. TiRex-2 activates 38.4M parameters |
| Streaming data | TiRex-2 | Recurrent xLSTM design with constant cost per new observation, per NXAI |
| Largest user base and tooling | Chronos-2 | Most downloaded time-series model on Hugging Face |
Licences: Open Weight Is Not Always Open Source
Permissive. Chronos-2, Toto 2.0, TiRex-2, TimesFM 2.5 and Sundial are Apache 2.0.
Restricted. The first TiRex uses the NXAI Community License, which adds commercial terms. TiRex-2 moved to Apache 2.0, with a paid Pro version for optimised streaming and fine-tuning. Granite PatchTST-FM r2 uses OpenMDW 1.0, a newer model licence, so read it before use.
Non-commercial. TimesFM 3.0 is under Google's TimesFM Non-Commercial License, a change from the Apache 2.0 terms of TimesFM 2.5. Moirai 2.0 is CC BY-NC 4.0, and Salesforce states the release is for research only.
The practical result is that the top-scoring model here cannot be used in a product. See the open-weight licence landscape for how these licences compare.
Caveats
- Training data overlap. Some models were pretrained on data that overlaps with the benchmark. The leaderboard labels each entry as zero-shot or pretrained and flags leakage. NXAI publishes separate decontaminated TiRex-2 checkpoints for this reason. The TiRex-2 row above is the zero-shot one. The variant trained with the GIFT-Eval pretraining set scores 0.678 and 0.467.
- Single models are not at the top of the board. A September 14, 2026 snapshot in the TW3Cast paper shows the first five places held by agent systems and routers that combine several models. The best single foundation model in that snapshot was a fine-tuned Toto 2.0 at position 23 of 130.
- The gaps are small. Eight models in the table sit between 0.667 and 0.719 on MASE. Results on your own data can reorder them.
- Baselines still matter. Test against a seasonal naive or statistical forecast before adopting any of these.
Brand Visibility Implications
Forecasting models run inside planning tools, monitoring products and data platforms. Their outputs are numbers in a dashboard, and no outside party can observe which model produced them. The visibility question is about selection. Data teams now ask an assistant which forecasting model to use, and the answer decides what gets tested. Developers of these models, and vendors that package them, have a stake in whether assistants describe the 2026 releases and their licences accurately. The same applies to the rest of this series, such as open-weight embedding models.
Methodology
This page is compiled from published sources, not from Presenc AI measurements. Parameter counts, licences and repository dates come from each model's Hugging Face card and the Hugging Face model API, read on October 1, 2026. Where a developer gives no release date, the month shown is the month the repository was created. Scores are quoted from the party that ran them and are labelled as such: developer model cards and papers, the GIFT-Eval leaderboard result files, the ViDoRe leaderboard as quoted on model cards, the openpi and Isaac GR00T repositories, and independent studies such as the Domyn guard model benchmark. Developer-reported scores use each developer's own test setup and are not directly comparable across rows. "Not reported" means we found no published score on that benchmark. Rankings in these categories change monthly. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks which models, libraries and vendors AI assistants recommend when users ask for a forecasting or time-series tool. Teams that build or sell these tools can see whether they are named and what the answers say about them.