CoRL 2026 · Conference on Robot Learning

Language as a Sensor: Calibrated Spatial Belief Estimation in 3D Scenes from Natural Language

Aryan Naveen1,2,* Jason Xinyu Liu2 Luca Carlone1,† Andreea Bobu2,†

1MIT Laboratory for Information & Decision Systems    2MIT Computer Science & Artificial Intelligence Laboratory

*Correspondence: aryannav@mit.edu    †Equal advising

A robot receives the utterance 'I left my backpack on the table'. The Language Sensor Model turns it into a multimodal spatial distribution over the tables in the scene graph, which VL-Map fuses with streaming visual observations until the belief concentrates on the true backpack location.
Language as a complementary sensing modality. Told “I left my backpack on the table,” the robot spreads belief over every plausible table, then sharpens it on the true location as visual observations arrive.

Overview

People constantly tell robots where things are, often about places the robot can’t see. We treat these utterances as a sensor measurement: a calibrated probability distribution over where the target is, fused with vision in a single Bayesian belief map.

Read the full abstract

Robots deployed in human-centric environments routinely receive natural-language descriptions of spatial information (“I left my backpack on the table”) that reference parts of the world beyond their perceptual field of view. Traditional metric-semantic mapping ignores this signal, while off-the-shelf multimodal models remain limited in 3D spatial reasoning and are not directly amenable to fusion with other sensor modalities. To convert language observations into a calibrated spatial distribution, we train a Language Sensor Model (LSM) that maps each utterance and its scene-graph context to a multimodal distribution, with mixture weights encoding referential ambiguity (e.g., which table) and component covariances encoding spatial uncertainty (e.g., where on the table the target lies). We then introduce VL-Map (Vision-Language Metric-Semantic Mapping), a probabilistic framework that treats these language predictions as stochastic observations and fuses them with onboard perception within a unified belief map. On the VLA-3D benchmark as well as on a real-world mobile robot, LSM is the only language predictor whose covariance estimates remain within the calibrated regime; fused into VL-Map, it leads to more accurate predictions of the target object location (~70% more probability mass on the true target compared to the strongest foundation-model baseline).

01

Language Sensor Model

Maps an utterance and scene graph to a calibrated Gaussian mixture over the target’s 3D position.

02

VL-Map

Fuses language and vision as independent observations in one voxel-level Bayesian belief.

03

Calibration pays off

The only calibrated language sensor we tested, with ~70% more belief on the true target than the best baseline.

Method

Language Sensor Model

“My backpack is on the table” is ambiguous in two ways, and LSM models each one explicitly.

Which table?Referential ambiguityAn LLM proposes and weights candidate anchors, which become the mixture weights.
Where on it?Spatial ambiguityA learned spatial transformer predicts a Gaussian per anchor, giving the covariances.
LSM pipeline: a hypothesis proposer enumerates anchor candidates from the prior scene graph and utterance; for each hypothesis the Language Spatial Transformer fuses object tokens with a BERT utterance encoding and decodes a Gaussian. Bottom: multimodal outputs for under, on, near, and between.
LSM pipeline. Top: hypothesis proposer → Language Spatial Transformer → Gaussian mixture. Bottom: outputs for under, on, near and between on the same scene.

One mode per plausible referent

Each proposed anchor contributes one mode, so belief stays on every plausible referent instead of collapsing to a single guess.

Office 1
Office 2
Living room 1
Living room 2
Lounge

VL-Map

Given the map, language and vision are conditionally independent, so the robot’s belief factors into one term per sensor.

belt(m)belief ∝ p(Z1:t | m, x1:t)vision · ∏j p(m | Lj, Mprior)language (LSM)

Implemented in Hydra as additive per-voxel log-odds. Vision updates every frame; each utterance adds one language update.

Results

~71%more belief on the true target than the best foundation-model baseline
1.83ANEES on val_seen (3 = calibrated), vs. 72.0 for Scaffolded-LLM
+4.34 natsinformation gain at the target on a real Boston Dynamics Spot

Static grounding

LSM is the most accurate and the only calibrated predictor. Post-hoc rescaling of LLM covariances doesn’t transfer to new scenes.

VLA-3D, unambiguous utterances with the ground-truth anchor (mean ± std). ANEES → 3 is calibrated; above 3 is overconfident.
Methodval_seenval_unseen
RMSE ↓NLL ↓ANEES → 3RMSE ↓NLL ↓ANEES → 3
Scaffolded-LLM1.78 ± 2.8735.82 ± 345.0672.01 ± 690.233.68 ± 5.5086.04 ± 370.05171.91 ± 740.27
+ global rescale1.78 ± 2.876.17 ± 19.924.16 ± 39.883.68 ± 5.509.33 ± 21.339.93 ± 42.77
+ per-axis rescale1.78 ± 2.875.94 ± 16.064.18 ± 32.183.68 ± 5.508.23 ± 15.928.23 ± 31.93
+ in-context1.77 ± 2.777.94 ± 38.5210.90 ± 77.093.45 ± 5.5923.75 ± 89.0742.28 ± 178.35
Scaffolded-VLM2.27 ± 1.6925.11 ± 49.4050.08 ± 99.062.79 ± 3.8639.92 ± 164.0479.57 ± 328.03
3D-ViSTA (fine-tuned)1.55 ± 1.642.33 ± 2.934.69 ± 10.663.66 ± 4.434.59 ± 4.059.08 ± 22.27
LSM (ours)0.71 ± 0.690.94 ± 1.671.83 ± 1.991.46 ± 1.684.17 ± 8.713.42 ± 16.71

As utterances get more ambiguous, baseline overconfidence compounds across mixture components. LSM stays in the calibrated band.

ANEES versus ambiguity level for each method, with the calibrated band from 2 to 4 shaded.
Calibration (ANEES) vs. ambiguity
RMSE versus ambiguity level for each method.
Accuracy (RMSE) vs. ambiguity

Closed-loop fusion

Overconfident language sensors make the fused belief worse than ignoring language. Only LSM improves on the uninformed prior throughout the rollout.

Information gain at the target surface over normalized episode progress for each method.
Information gain at the target (nats)
Mean target-object probability mass over normalized episode progress for each method.
Probability mass on the target
Rollout metrics, split by whether vision ever sees the target.
MethodMean IG (nats) ↑Mean prob. mass ↑Argmax dist. (m) ↓Success % ↑Miss % ↓
Target observed by vision (n = 44)
Vision-only−0.600.0065.520.079.5
LLM-E2E−6.220.0952.3522.752.3
Scaffolded-LLM−1.850.1302.2536.438.6
Scaffolded-VLM−2.260.0462.459.150.0
LSM (ours)+3.770.2231.4152.34.5
Target never directly observed (n = 10)
Vision-only−5.200.0075.470.080.0
LLM-E2E−1.150.1551.9740.040.0
Scaffolded-LLM−2.930.0793.2720.040.0
Scaffolded-VLM−4.890.0832.3610.060.0
LSM (ours)+1.650.2931.7570.030.0

Simulation rollout

“There is a cabinet that is over the counter.”

VL-Map (LSM)  finds the target
VL-Map (LLM)  confidently wrong

Real-world deployment on Spot

With a scene graph from an earlier mapping run, VL-Map puts belief on the target before Spot ever sees it. Vision alone has nothing until visual confirmation.

“I left my bag on the table.”

Vision-only
Language-only (LSM)
VL-Map (LSM + vision)  ours

BibTeX

@inproceedings{naveen2026language,
  title     = {Language as a Sensor: Calibrated Spatial Belief
               Estimation in 3D Scenes from Natural Language},
  author    = {Naveen, Aryan and Liu, Jason Xinyu and
               Carlone, Luca and Bobu, Andreea},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026}
}