OpenSLA

Sensor-Language-Action Models

Sensor
Language
Action
One model over sensor observations, individual context, and actions — evaluated across three healthcare settings.
One window, three questions

Ask what to do, what is happening, and why

The same model answers all three from one sensor window and the individual’s context.

OpenSLAMC-MED · 30-minute window

A patient’s 30-minute record from the emergency department. Ask the model any question about it.

The big idea

Sensor models stop at perception. Decisions don’t.

The point of sensing is not to describe the world. It is to decide what to do next.

01Sensing that stops before acting

Today’s sensor models are built to perceive: they recognize a state, or predict an outcome, and stop. Every decision that follows is modeled somewhere else, as its own task with its own closed label space. Nothing connects the evidence, the context, and the act.

02Language as the interface

Language can hold all of it at once: what is being asked, who the individual is, what the signals show, and what action means. Written in language, a bedside decision, a surgical intervention, and an insulin bolus become the same kind of problem — so one model can learn them all.

03One model that predicts, describes, and explains

OpenSLA is the first general framework to connect sensors, language, and actions. It decides whether to act, what kind of action, and which one; it describes what the signals show; and it explains the decision with the evidence behind it — all from a single model, grounded in the recording.

“Can sensor data, language, and action be modeled through a unified framework?”

What we found

Knows when to act — in every cohort.

Best action-necessity accuracy in all six cohorts. Best category in five.

  • OpenSLA
  • Strongest baseline

Balanced accuracy (%); 50 is chance. Baseline = strongest of the eight compared methods. Tables 2–4.

It describes the signal, not just the label.

Hover a phrase. The part of the recording it comes from lights up.

A 30-minute clinical window: ECG-II, plethysmography, respiration, respiratory rate, 1- and 5-minute HRV, heart rate, and perfusion.
Fig. 4

Ground truth

For heart-rate channel, the dominant finding is elevated heart-rate pattern; HR median 100, range 95–108, n=30. The cardiac waveform channel shows variable ECG waveform periodicity; rate proxy 95.2/min, IQR 0.40.

GPT-5.6-Luna zero-shot

At the decision time, the observed vital signs show sinus-rate cardiac activity with HR approximately 90 bpm.

OpenSLA

Across the recent monitored interval, pulse rate shows elevated heart-rate pattern; HR median 100, range 97–104, n=30. ECG Lead II: irregular Lead II rhythm proxy; rate proxy 95.2/min, IQR 0.40.

Results · state-estimation error, MAE ↓

1.17
Resp. rate MAE, clinical
next best 5.87
3.18
EtCO₂ MAE, operating room
next best 5.47
0.38
Carb amount MAE, CGM
next best 1.2

An action it never trained on lands beside its relatives.

Imipenem was never a training target, yet it lands among the antibiotics — and OpenSLA-H predicts it zero-shot at 71.4 balanced accuracy, against 61.9 for the strongest baseline.

(A) Zero-shot Imipenem prediction: balanced accuracy for OpenSLA, GPT-5.6-Luna, and OpenTSLM. (B) PCA of action representations with a zoomed inset, where the unseen action Imipenem lies inside the antibiotics cluster.
Fig. 5

Actions form a continuous space.

Interpolate between two embeddings, and the actions in between fall into place.

(A) Interpolating between the embeddings of two real samples (0.05 and 14.4 units of insulin) and retrieving the nearest real sample at each point gives 1.2 units at alpha 0.4 and 7.7 units at alpha 0.8. (B) The embedding of the current CGM window is read out into a future-state summary; mean absolute error over the next two hours is lowest for OpenSLA.
Fig. 6

Nearest real sample: 0.05 U endpoint

Language supervision is a first-order choice

  • Action-only training is the weakest variant on every endpoint. VitalDB necessity AUROC: 54.1 → 74.9 with captions and action conditioning.

  • Add action-evidence captions: necessity AUROC 81.6 → 83.4 on MOVER, 64.3 → 74.9 on VitalDB.

  • Drop action conditioning or the fusion decoder and you still beat action-only. The full model wins all four operating-room endpoints.

  • Direct LoRA beats Flamingo-style cross-attention and LLaVA-style projectors on all four AUROC endpoints. VitalDB necessity: 74.9 vs 68.9 vs 54.4.

  • Hierarchical Memory: 14.5× fewer sensor tokens in clinical windows, 4.7× in the operating room; forward passes 3.4× and 1.7× faster.

  • MIMIC-IV, zero fine-tuning: 74.4 vs 69.3 necessity balanced accuracy against the best adapted baseline.

Four variants side by side: OpenSLA with action-conditioned caption decoding, no action conditioning, no fusion decoder, and VLA-style action-only training, with a bar chart of necessity and category AUROC on VitalDB where OpenSLA is highest and VLA-style lowest.
Fig. 7
Three ways of connecting sensor tokens to the language model: direct LoRA adaptation as in OpenSLA, Flamingo-style cross-attention with a Perceiver resampler, and a LLaVA-style projector, with a bar chart of necessity and category AUROC on VitalDB where LoRA is highest.
Fig. 9
How it works

From raw recordings to grounded decisions

Caption for supervision. Train to act and to explain.

OpenSLA-B and OpenSLA-H architectures next to VLA-style, SensorLM, and supervised paradigms. OpenSLA-H adds Hierarchical Memory between the sensor encoder and the pretrained language model.
Fig. 3
  1. 01

    Encode every channel

    A text encoder for context, a frozen signal encoder per channel.

  2. 02

    Reason in the language model

    A LoRA-adapted LLM reads both. Action heads predict; an action-conditioned decoder explains.

  3. 03

    Compress with Hierarchical Memory

    Learned queries compress long recordings across resolutions; global queries keep the whole picture.

  4. 04

    The paradigms it is compared with

    VLA-style, SensorLM-style, and supervised — same data, same compute.

The language it learns from

Four caption facets, each grounded in the window

Hover a phrase. Its source lights up.

Individual context (60-year-old male, chief complaint abdominal pain and nausea, prior history) beside a 30-minute window of ECG-II, plethysmography, respiration, respiratory rate, heart rate, pain score, perfusion, and SpO2, with the recorded actions Cefepime, Azithromycin, and the target antiemetic.
Fig. 2

Observational what each channel shows

The lead II ECG showed an irregular-periodicity proxy, with an estimated rate of 100.2 beats/min. Respiratory rate was elevated, with a median of 26 breaths/min and a range of 23–30 across 31 measurements.

Relational how channels agree

Numerical respiratory-rate measurements and the respiration waveform both indicated rapid breathing.

Global findings in context

Overall findings suggested a pain and nausea symptom burden, supported by the presenting complaints, elevated heart rate, and a pain score of 5.

Action evidence for the recorded action

Findings that bear on pain or nausea management are high pain reported nausea provides symptom evidence to the antiemetic.

The benchmark

Three care settings, one action hierarchy

Sensor histories paired with individual context, structured actions, and evidence captions.

116k+
individuals
79
sensor modalities
60
action groups
7
datasets

Clinical care

Waveforms and vitals with clinical context. Medications, procedures, diagnostics.

MC-MEDMIMIC-IIIMIMIC-IV · external

Operating room

Perioperative waveforms and measurements, aligned with the actions taken in surgery.

MOVERVitalDB

Metabolic health

Glucose and basal insulin from daily life. The action is a bolus.

MetaboNetPEDAP
  1. NecessityAct or not?Yes
  2. CategoryWhat type of action?Antibiotics
  3. LabelWhich one?Imipenem
The paper

Perceive, reason, and act — through a common language

Language doesn’t just describe observations. It connects states to actions.

BibTeX
@misc{xu2026opensla,
  title   = {Sensor-Language-Action Models},
  author  = {Xu, Yuekai and Shuai, Zitao and Yang, Yuzhe},
  journal = {arXiv preprint arXiv:2610.08244},
  year    = {2026}
}