Lapras logo  Lapras: Latent Reasoning for Time Series Language Models

Paper | Code | Datasets

Lapras (Latent Post-trained Reasoning Across Series) is a post-training framework that equips time series language models (TSLMs) with latent reasoning. A model trained with Lapras reasons through a sequence of continuous thoughts in the joint time series–language space, producing text only for the final answer. It learns this through teacher–student self-distillation from reference chain-of-thought (CoT) traces.

This repository holds the 15 Lapras checkpoints reported in the paper: three TSLM backbones, each finetuned separately on five time series question answering tasks.

Checkpoints

Folder Backbone Finetuned from Size per checkpoint License
chatts/ ChatTS-8B (Qwen3-8B) bytedance-research/ChatTS-8B 16 GB Apache-2.0
opentslm/ OpenTSLM-Flamingo-1B (Llama 3.2 1B) OpenTSLM/llama-3.2-1b-sleep-flamingo 6.8 GB Llama 3.2 Community License
slip/ SLIP-1B (Llama 3.2 1B) leochen085/SLIP-Llama 3.3 GB Llama 3.2 Community License

Each backbone folder has one checkpoint per task: lapras_ecg, lapras_sleep, lapras_har, lapras_tsr, lapras_engine. A checkpoint answers questions of its own task only.

chatts/lapras_tsr/
├── config.json, generation_config.json, weights, tokenizer files
├── *.py                            # modeling code, loaded with trust_remote_code=True
├── lapras_projection.safetensors   # projection π that maps each continuous thought to the next input embedding
└── lapras_run_config.json          # inference settings (template, number of continuous thoughts, patch size)

The time series enters the language model differently in each backbone: as patch embeddings in the token sequence (ChatTS), through gated cross-attention over a Perceiver resampler (OpenTSLM-Flamingo), or through cross-attention in the last layers (SLIP).

How to use

The continuous-thought loop (K latent steps fed back through π before the answer is decoded) is implemented in the Lapras code, not in generate(). Run the checkpoints with its evaluation/evaluate.py, which reads lapras_run_config.json, so nothing has to be set by hand.

git clone https://github.com/yuc0805/Lapras.git && cd Lapras
# install as in the code README, then download into ckpt/ and dataset/:

hf download leochen085/Lapras --include "chatts/*" "opentslm/*" "slip/*" --local-dir ckpt  # all 15 checkpoints (128 GB)
hf download leochen085/Lapras --include "slip/*" --local-dir ckpt                           # one backbone
hf download leochen085/Lapras --include "slip/lapras_tsr/*" --local-dir ckpt                # one checkpoint

hf download leochen085/Lapras-Reasoning-Datasets --repo-type dataset --include "*.jsonl" --local-dir dataset

Evaluate one checkpoint on its test set with greedy decoding:

# single GPU
python evaluation/evaluate.py --ckpt ckpt/slip/lapras_tsr --test_file dataset/tsr/test.jsonl

# 4 GPUs
torchrun --nproc_per_node 4 evaluation/evaluate.py \
  --ckpt ckpt/chatts/lapras_tsr \
  --test_file dataset/tsr/test_with_ts_tags.jsonl \
  --batch_size 8 --max_new_tokens 512

ChatTS and OpenTSLM checkpoints read the *_with_ts_tags.jsonl files; SLIP reads the plain *.jsonl files. Predictions and metrics are written to <ckpt>/eval/ (predictions.jsonl, summary.json).

SLIP checkpoints fetch the configuration of the gated meta-llama/Llama-3.2-1B when loaded: accept its license and log in (hf auth login) first.

Tasks

The checkpoints were finetuned on the Lapras Reasoning Datasets: each example pairs a question and a multichannel time series with a reference reasoning trace that ends in Answer: <label>.

Task Description Test examples Channels × length Answer
ECG cardiological diagnosis 643 12 × 1000 yes / no
Sleep sleep-stage classification (SleepEDF) 923 1 × 1500 5 stages
HAR human activity recognition 8,222 3 × 128 8 activities
TSR counterfactual consequence prediction 4,094 1 × 128–1020 option letter + text (A–D)
Engine aero-engine fault diagnosis (EngineMT-QA) 193 33 × 600 option letter + text (A–D)

Results

Test accuracy and macro-F1 (%) of these checkpoints, as reported in Table 1 of the paper.

Backbone ECG Acc ECG F1 Sleep Acc Sleep F1 HAR Acc HAR F1 TSR Acc TSR F1 Engine Acc Engine F1 Avg Acc Avg F1
ChatTS-8B 85.23 85.17 88.19 79.49 77.43 73.53 69.32 69.32 35.75 35.10 71.18 68.52
OpenTSLM-Flamingo-1B 73.56 73.49 80.72 68.91 72.40 68.20 58.40 58.44 32.12 27.51 63.44 59.31
SLIP-1B 82.74 82.64 78.12 66.31 74.84 70.75 67.32 67.30 38.86 26.49 68.38 62.70

Compared with explicit CoT finetuning of the same backbones, Lapras raises the average F1 by 10.79 (ChatTS), 6.38 (OpenTSLM) and 5.07 (SLIP) points, while generating only the final answer.

License

The checkpoints follow the license of the model each was finetuned from (see LICENSE):

Llama 3.2 is licensed under the Llama 3.2 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.

Citation

@article{chen2026lapras,
  title   = {Lapras: Latent Reasoning for Time Series Language Models},
  author  = {Chen, Yuliang and Wu, Yu Yvonne and Langer, Patrick and Pillai, Arvind and Regmi, Sudarshan and
             Maritsch, Martin and Liu, Juncheng and Jakob, Robert and Kaar, Thomas and Griffin, Tess Z. and
             Marsch, Lisa and Heinz, Michael V. and Jacobson, Nicholas C. and Campbell, Andrew},
  journal = {arXiv preprint arXiv:2610.11111},
  year    = {2026}
}

Acknowledgements

We thank the authors of ChatTS, OpenTSLM and SLIP for releasing their models. Lapras follows the self-distillation objective of CODI.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train leochen085/Lapras

Paper for leochen085/Lapras