TemporalBench Leaderboard
This leaderboard presents offline evaluation results for agent configurations on the TemporalBench benchmark. It is a pure visualization and validation layer: no agents are executed here, and no LLM APIs are called.
Current entries: 7 agent configurations evaluated on all five V1.1 datasets. Entries measured only on the four V1.0 domains are not listed here โ switch to V1.0 to see them alongside these entries reduced to the same four domains. Rows are ranked by Overall T1/T3/T4 Acc, weighted by the corresponding MCQ question counts.
โ ๏ธ M5 was evaluated in a later round than the other four domains, after the DeepSeek API used for the original baselines was withdrawn. For the AgentScope, CAMEL, MetaGPT and TimeSeries Scientist deepseek-chat rows the overall score therefore averages across two API versions, and small gaps between them are not meaningful. See the About tab for details.
- "headers": [
- "Agent",
- "Agent Type",
- "Base Model",
- "Paper",
- "Evaluation Scope",
- "Domain Coverage",
- "All Domains Evaluated",
- "Overall T1 Acc",
- "Overall T2 Acc",
- "Overall T3 Acc",
- "Overall T4 Acc",
- "Overall T1/T3/T4 Acc",
- "Overall MCQ Acc",
- "Overall T2 MAE",
- "Overall T2 sMAPE",
- "Overall T4 MAE",
- "Overall T4 sMAPE",
- "MIMIC T2 OW-sMAPE",
- "MIMIC T2 OW-RMSSE",
- "MIMIC T4 OW-sMAPE",
- "MIMIC T4 OW-RMSSE"
- "data": [
- [
- "TimeClaw",
- "time-series-specific agent",
- "deepseek-v4-pro",
- "[arXiv:2606.05404](https://arxiv.org/abs/2606.05404)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4464,
- 0.3898,
- 0.4186,
- 0.429,
- 0.431,
- 0.4222,
- 48.2996,
- 60.4244,
- 44.2361,
- 58.0612,
- 9.4383,
- 161356.5029,
- 10.0052,
- 161228.102
- [
- "TimeCopilot",
- "time-series-specific agent",
- "deepseek-chat",
- "[arXiv:2509.00616](https://arxiv.org/abs/2509.00616)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4978,
- 0.2029,
- 0.3179,
- 0.2935,
- 0.3734,
- 0.3371,
- 0.1855,
- 0.7307,
- 0.194,
- 0.7742,
- 18.8553,
- 3.8524,
- 15.4094,
- 3.6373
- [
- "MetaGPT",
- "general agent",
- "deepseek-chat",
- "[arXiv:2308.00352](https://arxiv.org/abs/2308.00352)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4847,
- 0.2783,
- 0.2932,
- 0.3364,
- 0.3711,
- 0.3513,
- 59.1433,
- 28.628,
- 51.7442,
- 61.1881,
- 0.41,
- null,
- 0.41,
- null
- [
- "AgentScope",
- "general agent",
- "deepseek-chat",
- "[arXiv:2402.14034](https://arxiv.org/abs/2402.14034)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4792,
- 0.3029,
- 0.2793,
- 0.3406,
- 0.3649,
- 0.3517,
- 46.4439,
- 59.6662,
- 43.9194,
- 61.4402,
- 0.46,
- null,
- 0.39,
- null
- [
- "CAMEL",
- "general agent",
- "deepseek-chat",
- "[arXiv:2303.17760](https://arxiv.org/abs/2303.17760)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4398,
- 0.3068,
- 0.3115,
- 0.3446,
- 0.3648,
- 0.3525,
- 45.4158,
- 59.8788,
- 41.8789,
- 63.1245,
- 0.47,
- null,
- 0.46,
- null
- [
- "TimeCopilot",
- "time-series-specific agent",
- "deepseek-v4-pro",
- "[arXiv:2509.00616](https://arxiv.org/abs/2509.00616)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4465,
- 0.1593,
- 0.3825,
- 0.1194,
- 0.3333,
- 0.2962,
- 0.1947,
- 0.7251,
- 0.1701,
- 0.7571,
- 15.3372,
- 3.2549,
- 18.0584,
- 3.6529
- [
- "TimeSeries Scientist",
- "time-series-specific agent",
- "deepseek-chat",
- "[arXiv:2510.01538](https://arxiv.org/abs/2510.01538)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.3195,
- 0.2642,
- 0.2828,
- 0.2736,
- 0.293,
- 0.2868,
- 46.3897,
- 12.0962,
- 43.871,
- 12.0043,
- 12.34,
- null,
- 0.45,
- null
- [
- "metadata": null
Per-dataset results for all 13 entries, including those evaluated only on the four V1.0 domains โ Domain Coverage marks which is which. The newest evaluation domain, M5, is shown first. Use Select Columns to Display to add forecasting metrics or other domains.
- "headers": [
- "Agent",
- "Agent Type",
- "Base Model",
- "Paper",
- "Evaluation Scope",
- "Domain Coverage",
- "All Domains Evaluated",
- "Overall T1 Acc",
- "Overall T2 Acc",
- "Overall T3 Acc",
- "Overall T4 Acc",
- "FreshRetailNet T1 acc",
- "FreshRetailNet T2 acc",
- "FreshRetailNet T3 acc",
- "FreshRetailNet T4 acc",
- "PSML T1 acc",
- "PSML T2 acc",
- "PSML T3 acc",
- "PSML T4 acc",
- "Causal Chambers T1 acc",
- "Causal Chambers T2 acc",
- "Causal Chambers T3 acc",
- "Causal Chambers T4 acc",
- "MIMIC T1 acc",
- "MIMIC T2 acc",
- "MIMIC T3 acc",
- "MIMIC T4 acc",
- "FreshRetailNet T2 sMAPE",
- "FreshRetailNet T2 MAE",
- "PSML T2 sMAPE",
- "PSML T2 MAE",
- "Causal Chambers T2 MAE",
- "Causal Chambers T2 OW-RMSSE",
- "MIMIC T2 OW-sMAPE",
- "MIMIC T2 OW-RMSSE",
- "FreshRetailNet T4 sMAPE",
- "FreshRetailNet T4 MAE",
- "PSML T4 sMAPE",
- "PSML T4 MAE",
- "Causal Chambers T4 MAE",
- "Causal Chambers T4 OW-RMSSE",
- "MIMIC T4 OW-sMAPE",
- "MIMIC T4 OW-RMSSE",
- "M5 T1 acc",
- "M5 T2 acc",
- "M5 T3 acc",
- "M5 T4 acc",
- "M5 T2 sMAPE",
- "M5 T2 MAE",
- "M5 T4 sMAPE",
- "M5 T4 MAE"
- "data": [
- [
- "TimeClaw",
- "time-series-specific agent",
- "deepseek-v4-pro",
- "[arXiv:2606.05404](https://arxiv.org/abs/2606.05404)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4464,
- 0.3898,
- 0.4186,
- 0.429,
- 0.5625,
- 0.3106,
- 0.2102,
- 0.4394,
- 0.56,
- 0.2333,
- 0.364,
- 0.5,
- 0.22,
- 0.64,
- 0.488,
- 0.5133,
- 0.3245,
- 0.3333,
- 0.5686,
- 0.2837,
- 128.7568,
- 0.1134,
- 21.7273,
- 0.2112,
- 4.5199,
- null,
- 9.4383,
- 161356.5029,
- 123.9149,
- 0.104,
- 22.2906,
- 0.2253,
- 3.1434,
- null,
- 10.0052,
- 161228.102,
- 0.515,
- 0.42,
- 0.39,
- 0.4,
- 38.0957,
- 188.1663,
- 34.9564,
- 173.3398
- [
- "Single LLM",
- "single-LLM",
- "deepseek-chat",
- "โ",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC",
- "4 / 5",
- false,
- 0.4944,
- 0.3019,
- 0.351,
- 0.2879,
- 0.6818,
- 0.5682,
- 0.1685,
- 0.2652,
- 0.74,
- 0.2467,
- 0.348,
- 0.36,
- 0.1533,
- 0.1933,
- 0.444,
- 0.2933,
- 0.33,
- 0.2267,
- 0.3914,
- 0.2267,
- null,
- null,
- 0.27,
- 0.29,
- 1.97,
- 0,
- null,
- null,
- 0.97,
- 0.1,
- 0.28,
- 0.38,
- 2.53,
- 0,
- 0.48,
- null,
- null,
- null,
- null,
- null,
- null,
- null,
- null,
- null
- [
- "TimeCopilot",
- "time-series-specific agent",
- "deepseek-chat",
- "[arXiv:2509.00616](https://arxiv.org/abs/2509.00616)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4978,
- 0.2029,
- 0.3179,
- 0.2935,
- 0.5455,
- 0.2029,
- 0.1591,
- 0.381,
- 0.71,
- 0.2222,
- 0.288,
- 0.4912,
- 0.22,
- 0,
- 0.4,
- 0,
- 0.4574,
- 0.2366,
- 0.3542,
- 0.2727,
- 1.272,
- 0.1271,
- 0.2543,
- 0.2368,
- null,
- null,
- 18.8553,
- 3.8524,
- 1.3509,
- 0.1395,
- 0.2666,
- 0.2419,
- null,
- null,
- 15.4094,
- 3.6373,
- 0.49,
- 0.3611,
- 0.38,
- 0.3333,
- null,
- null,
- null,
- null
- [
- "MetaGPT",
- "general agent",
- "deepseek-chat",
- "[arXiv:2308.00352](https://arxiv.org/abs/2308.00352)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4847,
- 0.2783,
- 0.2932,
- 0.3364,
- 0.6989,
- 0.4167,
- 0.0511,
- 0.3712,
- 0.715,
- 0.1973,
- 0.28,
- 0.3067,
- 0.1533,
- 0.22,
- 0.464,
- 0.4,
- 0.3298,
- 0.227,
- 0.2579,
- 0.305,
- null,
- null,
- 25.15,
- 0.31,
- 2.33,
- 0.0024,
- 0.41,
- null,
- 125.21,
- 0.09,
- 29.31,
- 0.35,
- 2.7,
- 0.0028,
- 0.41,
- null,
- 0.46,
- 0.3467,
- 0.41,
- 0.3,
- 32.2509,
- 179.6086,
- 35.7077,
- 203.7171
- [
- "AgentScope",
- "general agent",
- "deepseek-chat",
- "[arXiv:2402.14034](https://arxiv.org/abs/2402.14034)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4792,
- 0.3029,
- 0.2793,
- 0.3406,
- 0.7045,
- 0.3485,
- 0.0568,
- 0.4242,
- 0.685,
- 0.2051,
- 0.248,
- 0.3133,
- 0.16,
- 0.4733,
- 0.464,
- 0.42,
- 0.3298,
- 0.227,
- 0.2571,
- 0.2482,
- 124.06,
- 0.1,
- 27.39,
- 0.29,
- 2.43,
- 0.0025,
- 0.46,
- null,
- 129.35,
- 0.07,
- 27.94,
- 0.36,
- 2.72,
- 0.0028,
- 0.39,
- null,
- 0.455,
- 0.26,
- 0.34,
- 0.3,
- 34.2596,
- 182.8507,
- 34.0857,
- 172.4051
- [
- "CAMEL",
- "general agent",
- "deepseek-chat",
- "[arXiv:2303.17760](https://arxiv.org/abs/2303.17760)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4398,
- 0.3068,
- 0.3115,
- 0.3446,
- 0.6705,
- 0.0833,
- 0.1989,
- 0.2955,
- 0.65,
- 0.1802,
- 0.252,
- 0.3067,
- 0.1267,
- 0.6533,
- 0.468,
- 0.48,
- 0.3138,
- 0.2482,
- 0.2561,
- 0.3121,
- 129.65,
- 0.11,
- 25.78,
- 0.28,
- 2.37,
- 0.0025,
- 0.47,
- null,
- 132.99,
- 0.22,
- 32.22,
- 0.39,
- 2.55,
- 0.0026,
- 0.46,
- null,
- 0.38,
- 0.34,
- 0.4,
- 0.32,
- 31.4413,
- 178.8021,
- 31.2733,
- 164.2514
- [
- "Single LLM",
- "single-LLM",
- "gpt-4o",
- "โ",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC",
- "4 / 5",
- false,
- 0.4972,
- 0.2984,
- 0.2924,
- 0.267,
- 0.6364,
- 0.5227,
- 0.0289,
- 0.1364,
- 0.675,
- 0.2067,
- 0.348,
- 0.36,
- 0.1333,
- 0.2733,
- 0.352,
- 0.26,
- 0.4681,
- 0.2128,
- 0.3661,
- 0.2979,
- 1.27,
- 0.12,
- 0.6,
- 0.61,
- 2.48,
- 0,
- 15.2,
- 0.55,
- 1.29,
- 0.34,
- 0.37,
- 0.44,
- 2.58,
- 0,
- 16.86,
- 0.63,
- null,
- null,
- null,
- null,
- null,
- null,
- null,
- null
- [
- "AgentScope",
- "general agent",
- "gpt-4o",
- "[arXiv:2402.14034](https://arxiv.org/abs/2402.14034)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC",
- "4 / 5",
- false,
- 0.4818,
- 0.2653,
- 0.2833,
- 0.2757,
- 0.625,
- 0.1212,
- 0.1364,
- 0.1894,
- 0.66,
- 0.2467,
- 0.272,
- 0.3533,
- 0.12,
- 0.46,
- 0.44,
- 0.32,
- 0.4468,
- 0.2128,
- 0.2395,
- 0.227,
- 126.27,
- 0.12,
- 37.38,
- 0.28,
- 2.76,
- 0.0026,
- 11.05,
- 0.43,
- 130.86,
- 0.2,
- 30.51,
- 0.35,
- 2.66,
- 0.0025,
- 12.02,
- 0.49,
- null,
- null,
- null,
- null,
- null,
- null,
- null,
- null
- [
- "CAMEL",
- "general agent",
- "gpt-4o",
- "[arXiv:2303.17760](https://arxiv.org/abs/2303.17760)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC",
- "4 / 5",
- false,
- 0.4944,
- 0.2618,
- 0.2558,
- 0.2792,
- 0.642,
- 0.0076,
- 0.0625,
- 0.3106,
- 0.685,
- 0.14,
- 0.184,
- 0.3067,
- 0.1,
- 0.66,
- 0.42,
- 0.2667,
- 0.4681,
- 0.2057,
- 0.3014,
- 0.234,
- 126.75,
- 0.13,
- 34.89,
- 0.43,
- 2.99,
- 0.0031,
- 12.02,
- 0.55,
- 128.18,
- 0.28,
- 35.78,
- 0.45,
- 2.5,
- 0.0026,
- 15.74,
- 0.59,
- null,
- null,
- null,
- null,
- null,
- null,
- null,
- null
- [
- "TimeCopilot",
- "time-series-specific agent",
- "deepseek-v4-pro",
- "[arXiv:2509.00616](https://arxiv.org/abs/2509.00616)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.4465,
- 0.1593,
- 0.3825,
- 0.1194,
- 0.5909,
- 0.2381,
- 0.1307,
- 0,
- 0.5204,
- 0.0588,
- 0.348,
- 0.1212,
- 0.1733,
- 0,
- 0.468,
- 0,
- 0.3564,
- 0.3,
- 0.52,
- 0.2708,
- 1.292,
- 0.142,
- 0.2263,
- 0.2411,
- null,
- null,
- 15.3372,
- 3.2549,
- 1.3791,
- 0.1422,
- 0.2097,
- 0.1946,
- null,
- null,
- 18.0584,
- 3.6529,
- 0.535,
- 0.2199,
- 0.37,
- 0.2033,
- null,
- null,
- null,
- null
- [
- "MetaGPT",
- "general agent",
- "gpt-4o",
- "[arXiv:2308.00352](https://arxiv.org/abs/2308.00352)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC",
- "4 / 5",
- false,
- 0.486,
- 0.2733,
- 0.2691,
- 0.2199,
- 0.625,
- 0.0909,
- 0.0511,
- 0.1439,
- 0.675,
- 0.2109,
- 0.22,
- 0.3133,
- 0.1067,
- 0.5933,
- 0.452,
- 0.16,
- 0.4574,
- 0.1702,
- 0.2897,
- 0.2553,
- 126.59,
- 0.13,
- 24.74,
- 0.34,
- 2.62,
- 0.0027,
- 14.11,
- 0.53,
- 127.22,
- 0.24,
- 43.47,
- 0.4,
- 2.76,
- 0.0029,
- 15.4,
- 0.63,
- null,
- null,
- null,
- null,
- null,
- null,
- null,
- null
- [
- "TimeSeries Scientist",
- "time-series-specific agent",
- "deepseek-chat",
- "[arXiv:2510.01538](https://arxiv.org/abs/2510.01538)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC, M5",
- "5 / 5",
- true,
- 0.3195,
- 0.2642,
- 0.2828,
- 0.2736,
- 0.3636,
- 0.5682,
- 0.3409,
- 0.5682,
- 0.385,
- 0.2667,
- 0.228,
- 0.2733,
- 0.2867,
- 0.0267,
- 0.296,
- 0.0267,
- 0.1543,
- 0.234,
- 0.2594,
- 0.234,
- 1.3,
- 0.17,
- 0.32,
- 0.34,
- 2.11,
- 0,
- 12.34,
- null,
- 1.25,
- 0.12,
- 0.27,
- 0.34,
- 2.51,
- 0,
- 0.45,
- null,
- 0.395,
- 0.26,
- 0.34,
- 0.3,
- 34.2596,
- 182.8507,
- 34.0857,
- 172.4051
- [
- "TimeSeries Scientist",
- "time-series-specific agent",
- "gpt-4o",
- "[arXiv:2510.01538](https://arxiv.org/abs/2510.01538)",
- "FreshRetailNet, PSML, Causal Chambers, MIMIC",
- "4 / 5",
- false,
- 0.2479,
- 0.2653,
- 0.2,
- 0.267,
- 0.3352,
- 0.5682,
- 0.0341,
- 0.5682,
- 0.28,
- 0.2667,
- 0.216,
- 0.2733,
- 0.2867,
- 0.0267,
- 0.216,
- 0.0267,
- 0.1011,
- 0.234,
- 0.2887,
- 0.234,
- 1.27,
- 0.35,
- 0.65,
- 1.53,
- 2.44,
- 0,
- 15.81,
- 0.52,
- 1.4,
- 0.51,
- 0.48,
- 0.84,
- 2.94,
- 0,
- 17.18,
- 0.64,
- null,
- null,
- null,
- null,
- null,
- null,
- null,
- null
- [
- "metadata": null
Upload submission files for manual review.
Required files:
results_on_dev_dataset.json: task-level metrics in leaderboard format.results_on_test_dataset.json: per-example test outputs with at leastid,tier,source_dataset,label, andoutput(required when the sample contains forecasting).
Please also include model architecture code and LLM/system details for verification.
Example record (JSON):
{
"agent_name": "AgentScope",
"agent_type": "general agent",
"base_model": "deepseek-chat",
"M5_T1_acc": 0.455,
"M5_T2_acc": 0.26,
"M5_T3_acc": 0.34,
"M5_T4_acc": 0.3,
"M5_T2_MAE": 182.85071428571425,
"M5_T2_sMAPE": 34.25956325164848,
"M5_T4_MAE": 172.40514285714286,
"M5_T4_sMAPE": 34.085729094579825,
"T1_acc": null,
"T2_acc": null,
"T3_acc": null,
"T4_acc": null,
"FreshRetailNet_T1_acc": 0.7045,
"FreshRetailNet_T2_acc": 0.3485,
"FreshRetailNet_T3_acc": 0.0568,
"FreshRetailNet_T4_acc": 0.4242,
"PSML_T1_acc": 0.685,
"PSML_T2_acc": 0.2051,
"PSML_T3_acc": 0.248,
"PSML_T4_acc": 0.3133,
"CausalChambers_T1_acc": 0.16,
"CausalChambers_T2_acc": 0.4733,
"CausalChambers_T3_acc": 0.464,
"CausalChambers_T4_acc": 0.42,
"MIMIC_T1_acc": 0.3298,
"MIMIC_T2_acc": 0.227,
"MIMIC_T3_acc": 0.2571,
"MIMIC_T4_acc": 0.2482,
"FreshRetailNet_T2_MAE": 0.1,
"FreshRetailNet_T2_sMAPE": 124.06,
"FreshRetailNet_T4_MAE": 0.07,
"FreshRetailNet_T4_sMAPE": 129.35,
"PSML_T2_MAE": 0.29,
"PSML_T2_sMAPE": 27.39,
"PSML_T4_MAE": 0.36,
"PSML_T4_sMAPE": 27.94,
"CausalChambers_T2_MAE": 2.43,
"CausalChambers_T2_OW_RMSSE": 0.00252,
"CausalChambers_T4_MAE": 2.72,
"CausalChambers_T4_OW_RMSSE": 0.00282,
"MIMIC_T2_OW_sMAPE": 0.46,
"MIMIC_T4_OW_sMAPE": 0.39
}
The paper describing this benchmark is TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks (https://arxiv.org/abs/2602.13272). We also maintain a public leaderboard and welcome submissions from state-of-the-art models: https://huggingface.co/spaces/Melady/TemporalBench_Leaderboard
What this leaderboard shows
- One row per evaluated agent configuration, with a Paper link to the source
publication for each agent framework (
Single LLMis a raw-model baseline and has none) - Task-family MCQ metrics for TemporalBench (T1โT4)
- Forecasting metrics for T2/T4 (sMAPE, MAE) and MIMIC OW metrics when provided
- Dataset-level results for: FreshRetailNet, PSML, Causal Chambers, MIMIC, and M5
Data requirements
Results are loaded from a local JSON or CSV file. Each record must include:
- Identity fields:
agent_name,agent_type,base_model - Required metrics:
T1_acc,T2_acc,T3_acc,T4_acc(computed overall) - Optional metrics:
- Overall forecasting:
T2_sMAPE,T2_MAE,T4_sMAPE,T4_MAE - MIMIC overall OW:
MIMIC_T2_OW_sMAPE,MIMIC_T2_OW_RMSSE,MIMIC_T4_OW_sMAPE,MIMIC_T4_OW_RMSSE - Dataset-level metrics:
<Dataset>_T{1..4}_accand forecasting metrics per dataset
- Overall forecasting:
Overall computation
Overall T1โT4 accuracy and T2/T4 forecasting metrics are computed as weighted averages from dataset-level results using question/series counts, over the domains of the selected benchmark version.
Each version's leaderboard ranks only entries evaluated on that version's full domain set, so a ranking always compares entries measured on the same domains. V1.1 therefore lists only entries with results for all five datasets. V1.0 is backward compatible: it lists entries that predate M5 together with V1.1 entries reduced to their four-domain subset, recomputed and re-ranked. The By Domain tab is unfiltered and reports per-dataset results for every entry, with Domain Coverage marking which entries are which.
M5 evaluation timeline and API availability
M5 was added to TemporalBench after the original four domains (FreshRetailNet, PSML, Causal Chambers, MIMIC) had already been evaluated. By the time M5 was run, the DeepSeek API used for the original baseline runs was no longer served, so those baselines could not be re-run end to end under a single API version.
Four deepseek-chat rows therefore combine two evaluation rounds:
- AgentScope, CAMEL, MetaGPT and TimeSeries Scientist โ the four original domains are the values reported in the paper; M5 was evaluated separately, later, against the DeepSeek API available at that time.
Every other row comes from a single evaluation round. TimeCopilot (deepseek-chat),
TimeClaw (deepseek-v4-pro) and TimeCopilot (deepseek-v4-pro) were each run across all
five domains in one pass. The gpt-4o rows and Single LLM (deepseek-chat) predate M5 and
carry no M5 results, so they appear in the V1.0 leaderboard and the By Domain tab but
not in V1.1.
How to read this: for the four rows above, Overall T1/T3/T4 Acc averages across that seam,
so part of the row reflects one API version and part reflects another. Those rows currently
sit within roughly 0.01 accuracy of their neighbours, which is small relative to the
variation a model or API version change can introduce โ differences of that size between
them should not be read as stable differences between the systems. Per-domain values in the
By Domain tab, and comparisons between rows that share a single evaluation round, are
unaffected.
The Benchmark version switch on the Leaderboard tab removes this seam entirely: selecting V1.0 restricts every entry to the four original domains, recomputes the overall metrics from those domains only, and re-ranks the table, so all entries are compared on the same domain set.
Submission workflow
Uploads are stored locally for manual review.
For a valid submission, please provide two files:
results_on_dev_dataset.json- This follows the leaderboard metrics format.
- It should include task-level metrics only (e.g., T1-T4 and forecasting metrics).
results_on_test_dataset.json- This should include per-example outputs on the test split.
- For each example, include at least:
idtiersource_datasetlabeloutput(required when the example contains a forecasting task)
We also strongly encourage including model and system metadata, such as:
- model architecture code
- LLM(s) used
- key implementation details needed for result verification
Approved submissions should then be merged into the main results file to appear on the leaderboard.
Data access
The dataset is available at:
https://huggingface.co/datasets/Melady/TemporalBench
It includes all test tasks and a forecast_metrics_utils.py file that documents the
standard metric computation utilities.
Citation
Copy the following snippet to cite these results
@misc{weng2026temporalbenchbenchmarkevaluatingllmbased,
title={TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks},
author={Muyan Weng and Defu Cao and Wei Yang and Yashaswi Sharma and Yan Liu},
year={2026},
eprint={2602.13272},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2602.13272},
}