Large Language Models (LLMs) are increasingly being applied to time-series forecasting, giving rise to a class of models referred to as time-series LLMs. While these models achieve competitive predictive accuracy, their reliability and structural consistency remain insufficiently understood. These models may produce forecasts that are numerically accurate on average yet statistically or temporally implausible, a phenomenon referred to as hallucination. Unlike conventional forecasting errors, hallucinations represent deviations from underlying temporal dynamics that exceed expected volatility patterns. This paper presents a systematic investigation of hallucination in time-series LLM forecasting. We introduce a quantitative evaluation framework that complements traditional regression metrics with two reliability-oriented measures: perplexity (PP), which reflects predictive uncertainty, and hallucination rate (HR), which measures statistically significant deviations from ground truth. Experiments on widely used benchmark datasets (Electricity and ETT variants) reveal a critical trade-off between forecasting accuracy and reliability. In several settings, improvements in average error metrics do not correspond to improved reliability; models can maintain low mean absolute error (MAE) while exhibiting high HR. To mitigate this issue, we evaluate two strategies: data-centric preprocessing, whose effectiveness depends on dataset characteristics, and structured tokenization, which consistently reduces hallucination across the evaluated datasets. Sensitivity analysis over quantile thresholds confirms the robustness of hallucination trends and model rankings. These results demonstrate that conventional regression metrics alone are insufficient for evaluating time-series LLMs and highlight the need for reliability-focused diagnostics when deploying LLM-based forecasting systems in high-stakes domains.



This work was supported by Yayasan Universiti Teknologi PETRONAS (YUTP) under the Pre-commercialisation Research Grant (YUTP-PRG), Grant Title: Integration of Metocean Data for Intelligent Operations and Automation using Large Language Models (LLMs) (Cost Centre: 015PBC-035).
Department of Computing, Faculty of Management, Science and Computing, Universiti Teknologi PETRONAS, Seri Iskandar, 32610, Perak, Malaysia