TSFM-RouterBench

A controlled benchmark for routing frozen time-series foundation models.

— entries · code version — · built — · —

Provisional results. Scores are official leaderboard metrics (per-config official metric over Seasonal Naive, geometric mean across the 97 GIFT-Eval configs) computed from frozen, sha256-stamped routing decisions. They are pre-audit: the two-author cross audit is still in progress, and the Equal Weight combiner awaits its official scorer pass.

About

Router comparisons usually confound three things at once: which candidate models were available, how much history the router was allowed to see, and at what granularity it decided. This benchmark holds all three fixed and states them on every row. Candidate forecasts are computed once and frozen into a public bank, so evaluation only reads files; an information firewall hands the router evidence windows and nothing else, and the decisions are hashed to a manifest before any test label becomes readable. Scores are then decomposed into headroom H — what the pool made possible — and gain G — what the router actually took — so that a strong pool is never mistaken for a strong router.

Anyone can add a row. Run your router locally against the frozen bank, check the package with validate_submission.py, open a pull request, and a maintainer reviews it before it is merged and listed. There is no CLA, and a submission is a few hundred kilobytes of reviewable text rather than a pile of forecast arrays. See the submission guide, the directory schema, or the repository.

Experimental conditions

Results are comparable only within one setting of these controls. Changing any of them changes what the numbers mean. One choice per row; the options are the full protocol grid, so a setting with nothing in it yet simply shows an empty table.

Benchmark
Pool size
Validation
Train windows

Method filters

These rows are multi-select, and a method may carry more than one mechanism tag. Decision labels are not mutually exclusive either: a hybrid method answers to both Hard and Soft, and still occupies a single row in the table.

Decision
Granularity
Mechanism
Source

Overall table

— Entries
— Benchmarks covered
— Pool sizes covered
— Best HR (hard)

Sorted by the internal scale-normalized pinball loss, ascending, by default; click any header to change the sort, or any row to expand its provenance. This diagnostic loss is not the backend-native official CRPS. Formal releases report official MASE/CRPS separately. HR (headroom realized, G / H) is reported for hard-selection methods only: the member-level oracle bounds the action space of picking one member, so a soft combination can legitimately beat it and the ratio stops meaning "share of the available headroom". Combiners show — in that column and are compared on loss and G instead.

Method Mechanism Decision Granularity Evidence LB-CRPS LB-MASE G HR Source

How to read a row

Evidence
The causal window budget the router was allowed to see before deciding: val_only is the validation window alone, val+Bk adds the k most recent training backtest windows per series, Bk uses those windows without validation.
Mechanism
What the method decides on. A method can carry several tags at once — a gate over engineered features that falls back on backtest losses is feature+backtest, and it answers to either filter.
G
Router gain: how much loss the router removed relative to the best fixed single member of the same pool under the same evidence.
H
Headroom: how much a per-task oracle over the same pool would have removed. It is a property of the pool, not of the method.
HR
Headroom realized, G / H. Near-zero headroom makes the ratio meaningless, so it is withheld rather than divided out.
Source
reproduced — results produced and verified by the benchmark maintainers on their own hardware. community — results submitted by external researchers and listed as submitted after maintainer review; the numbers were not re-run by the maintainers. Submissions still awaiting review sit in the repository and are not shown.

Protocol at a glance

Pool track

Controls which candidate models the router could choose between. The benchmark evaluates every eligible composition at each pool size; a nested representative chain may be shown only as a visualization and never substitutes for composition dispersion.

Because headroom is a property of the pool, every score is reported next to the headroom its own pool allowed.

Evidence track

Controls how much history the router saw before deciding: the validation window closest to the test period, plus the k most recent training backtest windows per series, taken in causal order.

The grid uses validation alone, and validation or train-only conditions with caps of 3, 10, 100, and 300 recent training windows per series. A cap is a maximum: shorter series use every legal window they actually have, and the gap is reported per task. val_only is the pre-registered reference point; a conclusion is only claimed when it survives the other settings.

Granularity track

Controls how often a decision may change — once per task, once per series, or once per forecast window. Methods run at their native granularity rather than being reimplemented.

Each is scored against an oracle at the matching granularity, so a finer decision space is never credited with headroom it did not have.