Introduction
Prompt pipelines look weak on time-series classification. Same model internals, read with a linear probe, tell a different story.
Figure 1. Main idea: prompt-based generation underestimates representational capacity revealed by linear probing.
Paper
Prompt-based prediction pipelines suggest that large language models (LLMs) perform poorly on time-series classification, raising doubts about whether they encode meaningful temporal structure. We show that this conclusion reflects limitations of prompt-based generation rather than the model's representational capacity by directly comparing prompt outputs with linear probes over the same internal representations while controlling for important confounding factors. While prompting performs near chance, linear probes improve average F1 from 0.15-0.26 to 0.61-0.67, often matching or exceeding specialized state-of-the-art (SoTA) time-series models. Layer-wise analyses further show that class-discriminative time-series information emerges in early transformer layers and is amplified with multimodal inputs and the presence of pretrained weights. Together, these results demonstrate a systematic mismatch between what LLMs internally represent and what prompt-based prediction pipelines reveal.
Setup
Digits (d), visualizations (v), or both (dv). Same prompt used for generation (PBC) and representation extraction (RP). No fine-tuning, no extra TS embedders.
Figure 2. Overview of prompting vs. probing paradigms. Raw time series are transformed into digit-based text representations and visualizations which are incorporated into prompts for direct prediction or probing.
Results
On dv, probes jump F1 from ~0.15-0.26 to ~0.61-0.67 and compete with Attend / TS2Vec / Moment. Prompting stays near chance.
| Method | CTU | EMG | HAD | HAR | RWC | TEE | Avg |
|---|---|---|---|---|---|---|---|
| Attend | .848 | .985 | .732 | .933 | .726 | .524 | .791 |
| TS2Vec | .631 | .933 | .667 | .865 | .672 | .794 | .761 |
| Moment | .660 | 1.00 | .671 | .878 | .778 | .564 | .758 |
| OneFitsAll | .587 | .296 | .681 | .885 | .790 | .475 | .615 |
| InstructTime | .627 | .167 | .278 | .527 | .434 | .235 | .378 |
| Llama RP | .640 | 1.00 | .406 | .723 | .672 | .223 | .611 |
| Llama PBC | .356 | .175 | .033 | .111 | .190 | .036 | .150 |
| Mistral RP | .684 | .933 | .438 | .730 | .645 | .462 | .649 |
| Mistral PBC | .548 | .393 | .021 | .127 | .190 | .288 | .261 |
| Qwen RP | .676 | 1.00 | .374 | .737 | .656 | .592 | .672 |
| Qwen PBC | .324 | .413 | .016 | .174 | .222 | .304 | .242 |
Table 1. Macro F1 across datasets. Prompt (PBC) and Probe (RP) results for the dv modality. Bold: probe averages.
Analysis
Class-discriminative structure appears early. Multimodal inputs and pretrained weights amplify deeper layers.
Figure 3. Layer-wise probe macro F1 for Mistral and Qwen across modalities.
Figure 4. Layer-wise macro F1 of probes for Qwen versus the same model with randomly initialized weights.
Reliability
Prompt wording and sampling swing PBC scores. Visual-only prompting is a bit more stable, but still far below probes.
Figure 5. Prompt-based macro F1 distributions by input modalities, showing a bit more stable performance for visual-only prompting.
Conclusion
@misc{schumacher2026promptingunderestimatesllmcapability,
title={Prompting Underestimates LLM Capability for Time Series Classification},
author={Dan Schumacher and Erfan Nourbakhsh and Rocky Slavin and Anthony Rios},
year={2026},
eprint={2601.03464},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2601.03464},
}