LAGEN hourglass logo

LAGEN

Are Vision-Language Agents Latency Aware?

Training agents for a world that keeps moving

* Equal contribution; author order is interchangeable.
† Work done during internship at Northwestern.

1 Northwestern University2 University of Massachusetts Amherst 3 Stanford University4 The University of Texas at Austin5 NVIDIA

Timing changes the outcome

Both runs start from the same state and replay the same actions. Use the slider to delay the actions on the right.

Pausing the world hides the cost of latency

While an agent computes, the world changes. Its action then reaches a different state from the one it observed. On Flappy Bird and Demon Attack, a two-frame delay leaves agents with just 1.2–7.1% of their zero-latency return.

OpenVLAπ₀.₅GR00T Single action Action chunk (horizon 8)

Flappy Bird

Eval latency (frames) Mean return 0 100 200 300 400 500 0 1 2 3 4 Zoom: latency > 0 0 10 20 1 2 3 4

Demon Attack

Eval latency (frames) 0 1000 2000 3000 4000 5000 6000 7000 0 2 4 Zoom: latency > 0 100 200 300 2 4

Deadly Corridor

Eval latency (frames) 0 100 200 300 400 500 600 700 0 1 2 3 4 5 Zoom: latency > 0 80 160 240 1 3 5

InterceptGrabFast

Eval latency (frames) Success rate (%) 0 10 20 30 40 50 60 70 80 90 100 0 1 2 3 4

Focus a chart and use the left and right arrow keys to inspect latency columns. Press Escape to dismiss values. On touch screens, tap a column.

Each point uses a fixed action delay. The game plots show mean return. InterceptGrabFast shows success rate over 200 episodes.

Reducing latency is not enough. Train with it

Reduce latency

Accelerate inference. Predict new actions while executing the current ones.

Train with latency (ours)

Reproduce deployment latency in simulation. Train agents to act under these delays.

Training with deployment latency improves real-time performance across all 36 combinations of tasks and architectures.

OpenVLAπ₀.₅GR00T
Before training After training Zero latency (100%)
GameEmbodied 0% 25% 50% 75% 100% 125% 150% 175% Relative performance (%) Demon Attack Asterix Atlantis AirRaid Deadly Corridor Flappy Bird Ant Hopper Inverted Pendulum Walker2D Balance Intercept Grab Fast
0% 25% 50% 75% 100% 125% 150% 175% Relative performance (%) Demon Attack Asterix
0% 25% 50% 75% 100% 125% 150% 175% Relative performance (%) Atlantis AirRaid
0% 25% 50% 75% 100% 125% 150% 175% Relative performance (%) Deadly Corridor Flappy Bird
0% 25% 50% 75% 100% 125% 150% 175% Relative performance (%) Ant Hopper
0% 25% 50% 75% 100% 125% 150% 175% Relative performance (%) Inverted Pendulum Walker2D
0% 25% 50% 75% 100% 125% 150% 175% Relative performance (%) Balance Intercept Grab Fast

Hover or tap to inspect values. Focus a chart and use the arrow keys to move between readings. Press Escape to close.

RTX 3090 latency profiles · 12 tasks · 3 VLA architectures. The hatched region shows performance before latency training. The full bar shows performance after training. Both use the original policy’s zero-latency return as 100%.
All 36 results

The table reports mean episode returns. Zero/Zero uses a policy trained and evaluated without latency. Zero/Real evaluates the same policy in real time. Target/Real evaluates a policy trained with the deployment latency profile in real time.

TaskArchitectureZero/ZeroZero/RealTarget/Real
Demon AttackOpenVLA2362.2584.41452.8
Demon Attackπ₀.₅2350.9112.3907.3
Demon AttackGR00T2370.9231.91203.1
AsterixOpenVLA5803.5374.53653.5
Asterixπ₀.₅5758.5143.51015.0
AsterixGR00T5714.0182.01440.0
AtlantisOpenVLA39560.018588.039229.0
Atlantisπ₀.₅39615.04369.026447.0
AtlantisGR00T39662.08167.028488.0
AirRaidOpenVLA2000.0878.82353.8
AirRaidπ₀.₅1943.81386.33227.5
AirRaidGR00T2101.3733.81586.3
Deadly CorridorOpenVLA2098.7742.42091.2
Deadly Corridorπ₀.₅2104.1140.91450.8
Deadly CorridorGR00T2100.5341.82090.8
Flappy BirdOpenVLA439.119.2373.8
Flappy Birdπ₀.₅408.75.6344.5
Flappy BirdGR00T428.56.6414.5
AntOpenVLA5565.3194.31853.3
Antπ₀.₅5782.5131.61774.7
AntGR00T6222.6165.41956.0
HopperOpenVLA1419.967.9424.6
Hopperπ₀.₅3482.58.4235.2
HopperGR00T872.88.6242.6
Inverted PendulumOpenVLA1000.019.8586.9
Inverted Pendulumπ₀.₅1000.020.1830.4
Inverted PendulumGR00T1000.015.61000.0
Walker2DOpenVLA983.9103.7857.4
Walker2Dπ₀.₅3631.481.0371.1
Walker2DGR00T3528.063.2852.9
BalanceOpenVLA275.650.7114.6
Balanceπ₀.₅328.646.872.3
BalanceGR00T267.545.390.4
Intercept Grab FastOpenVLA36.118.625.8
Intercept Grab Fastπ₀.₅36.88.117.6
Intercept Grab FastGR00T33.514.325.2

Reproducing deployment latency in simulation

Simulation predicts deployment performance

We reproduce deployment latency in simulation to predict how agents perform. Temporal replay reduces normalized return error from 8.93 to 3.81 percentage points compared with a fixed mean delay. This comparison covers five GPU types, three tasks, and two models.

Flappy BirdInverted PendulumInterceptGrabFast
RTX 3090RTX 4090RTX 5090L40SA100
OpenVLA
Fixed
0 0 25 25 50 50 75 75 100 100 Real mean return (%) Simulated mean return (%)
Normal
0 0 25 25 50 50 75 75 100 100 Real mean return (%) Simulated mean return (%)
IID
0 0 25 25 50 50 75 75 100 100 Real mean return (%) Simulated mean return (%)
Temporal
0 0 25 25 50 50 75 75 100 100 Real mean return (%) Simulated mean return (%)
GR00T
Fixed
0 0 25 25 50 50 75 75 100 100 Real mean return (%) Simulated mean return (%)
Normal
0 0 25 25 50 50 75 75 100 100 Real mean return (%) Simulated mean return (%)
IID
0 0 25 25 50 50 75 75 100 100 Real mean return (%) Simulated mean return (%)
Temporal
0 0 25 25 50 50 75 75 100 100 Real mean return (%) Simulated mean return (%)

Hover or tap to inspect values. Focus a chart and use the arrow keys to move between readings. Press Escape to close.

Both axes show return relative to the same zero-latency baseline (100%). The diagonal marks equal simulated and measured returns. Sim − real shows their difference in percentage points.
How we test the simulation

We evaluate OpenVLA and GR00T on three tasks: Flappy Bird, Inverted Pendulum, and InterceptGrabFast. We use five GPU types: RTX 3090, RTX 4090, RTX 5090, L40S, and A100.

Each combination uses 25 runs to fit the latency model and 25 separate runs to evaluate it. Each run contains two episodes. The simulation models both action delays and inference capacity.

Overall normalized return error, in percentage points: Fixed 8.93, Normal 6.55, IID 4.58, and Temporal 3.81. See paper §3.4.

Slow requests occur in bursts

Two deployments can have the same average latency but different sequences of delays. Temporal profiles capture how often different delays occur and how slow requests cluster over time.

Action latency (ms)

5075100 010k20k Request index

Density (ms⁻¹)

0.25.5 5075100 Latency (ms)

Action latency (ms)

5075100 010k20k Request index

Density (ms⁻¹)

0.25.5 5075100 Latency (ms)

Action latency (ms)

5075100 010k20k Request index

Density (ms⁻¹)

0.25.5 5075100 Latency (ms)

Action latency (ms)

5075100 010k20k Request index

Density (ms⁻¹)

0.25.5 5075100 Latency (ms)

Action latency (ms)

5075100 010k20k Request index

Density (ms⁻¹)

0.25.5 5075100 Latency (ms)
● Worker 0■ Worker 1Measured distribution

Hover or focus to highlight. Click or press Enter to keep a selection; repeat to clear it. Use arrow keys to move between choices. Escape clears the selection.

Measured latency and sequences generated by four fitted models. Select a latency range to highlight matching requests across all five sequences. The displayed range is 45–100 ms. Distributions use the full sequences.

Training agents with deployment latency

We replay deployment latency in simulation while the world keeps moving. A teacher learns to act under these delays through reinforcement learning. We then fine-tune the VLA using the teacher’s action labels.

LAGEN pipeline from the paper: profile deployment timing using Fixed, Normal, IID or Temporal generators; train a latency teacher through fresh reinforcement-learning interactions; distill teacher demonstrations into a VLA with optional DAgger; deploy the VLA with delayed-action execution.
The agent interacts with the simulator under delays sampled from the latency profile. ActionReady marks when a predicted action becomes available. The curves show schematic latency profiles.

Key Findings

1. Profile-trained policies achieve higher returns

Training with the full latency profile exposes agents to varying delays and consecutive slow requests. Policies trained with the full profile achieve higher returns than those trained at the mean on three of four tasks. Flappy Bird returns are similar.

OpenVLA · one action per request · mean episode return ± standard deviation
TrainingFlappy BirdDeadly CorridorAntInterceptGrabFast
Fixed mean384.82± 116.791620.80± 913.621453.84± 693.733.54± 7.07
Full profile373.78± 126.832091.23± 579.611853.33± 508.6918.20± 15.34
Profile gain-2.9%+29.0%+27.5%+413.4%

2. Transfer is stronger from high to low latency

Transfer is stronger from longer delays to shorter delays than in the reverse direction. Flappy Bird performs best near its training delay. Scores use a policy trained at the evaluation delay as the reference.

Relative to matched-delay policy

0%25%50%75%100%
Flappy Bird
Eval latency (raw frames)
4 1.1% 1.3% 1.1% 2.7% 100.0%
3 1.5% 1.4% 1.3% 100.0% 7.0%
2 1.5% 2.2% 100.0% 3.7% 1.8%
1 4.1% 100.0% 1.6% 1.6% 1.4%
0 100.0% 19.7% 1.8% 1.5% 1.4%
01234
Train latency (raw frames)
Demon Attack
Eval latency (raw frames)
32 4.8% 8.1% 37.8% 87.5% 100.0%
24 3.3% 6.0% 24.3% 100.0% 65.2%
16 4.6% 27.4% 100.0% 66.5% 51.8%
8 4.3% 100.0% 84.9% 64.6% 32.6%
0 100.0% 64.0% 36.5% 24.6% 23.0%
08162432
Train latency (raw frames)
VizDoom Deadly Corridor
Eval latency (raw frames)
32 36.2% -0.2% 65.4% 52.9% 100.0%
24 90.3% 76.0% 323.2% 100.0% 278.2%
16 65.2% 44.1% 100.0% 121.7% 122.9%
8 45.3% 100.0% 74.2% 79.8% 68.4%
0 100.0% 26.8% 50.0% 50.1% 25.9%
08162432
Train latency (raw frames)
InterceptGrabFast
Eval latency (raw frames)
4 13.2% 19.3% 33.3% 80.7% 100.0%
3 33.8% 50.8% 84.6% 100.0% 89.2%
2 66.9% 91.9% 100.0% 91.9% 67.4%
1 60.1% 100.0% 100.0% 80.9% 64.9%
0 100.0% 96.4% 96.4% 77.9% 62.6%
01234
Train latency (raw frames)

Hover or tap to inspect values. Focus a chart and use the arrow keys to move between readings. Press Escape to close.

Arrows show training → evaluation latency. High → low latency is the mean of the ten cell percentages in the lower-right triangle. Low → high latency is the corresponding mean for the upper-left triangle.

3. Latency training transfers to other tasks

We train on Demon Attack with mixed latencies and evaluate the policy on other games. It achieves higher returns under delay on Atlantis and AirRaid than a policy trained without latency. On Asterix, it outperforms this baseline at two frames and underperforms at four.

Zero-latency trainMixed-latency train
Demon Attack
0 2k 4k 6k 0 2 4 Mean return Eval latency (raw frames)
Asterix
0 100 200 300 0 2 4 Mean return Eval latency (raw frames)
Atlantis
0 1k 2k 3k 4k 5k 0 2 4 Mean return Eval latency (raw frames)
AirRaid
0 100 200 300 400 0 2 4 Mean return Eval latency (raw frames)

Hover or tap to inspect values. Focus a chart and use the arrow keys to move between readings. Press Escape to close.

Both policies train on Demon Attack. Each point shows mean return over 20 episodes. Returns use each task’s native scale.

4. Explicit latency information improves control

The same observation can require different actions at different latencies. Stating the latency in the prompt lets the agent distinguish these cases. On Flappy Bird, mean return rises from below 20 to 268–434 across delays of 0–4 frames. On Demon Attack, mean return improves at delays of 0, 1, 2, and 4 frames; the means are similar at three frames.

Mean return ± 95% confidence interval · 20 episodes per condition
TaskDelay in prompt0 frames1 frame2 frames3 frames4 frames
Flappy BirdYes433.8± 15.7434.0± 22.2341.3± 69.2290.3± 81.2267.6± 69.0
Flappy BirdNo18.8± 7.314.3± 7.39.4± 3.16.5± 1.64.5± 0.4
Demon AttackYes7409.0± 302.66558.0± 509.75551.2± 474.24038.0± 896.43630.0± 717.4
Demon AttackNo3814.8± 947.94392.5± 851.84478.8± 516.94091.2± 877.51995.2± 705.5

The prompt states the delay shown in each column.

5. Visual history improves control

Visual history shows how the scene changes over time. Ghost trails improve Flappy Bird performance, KV memory gives the largest gain on Demon Attack, and a frame grid performs best on Deadly Corridor.

Return (% of single-frame baseline)

Flappy Bird

3 frames of delay

50 75 100 125 150
Demon Attack

6 frames of delay

50 75 100 125 150
Deadly Corridor

6 frames of delay

50 75 100 125 150 N/A

Hover or focus to highlight. Click or press Enter to keep a selection; repeat to clear it. Use arrow keys to move between choices. Escape clears the selection.

100% marks the single-frame baseline at the same delay. Delays are three frames on Flappy Bird and six on Demon Attack and Deadly Corridor. Error bars show episode-level 95% intervals from one training run. OpenVLA uses 100 evaluation episodes per condition, and WanOFT uses 20. Ghost trail was not evaluated on Deadly Corridor.

Citation

@misc{chafekar2026lagen,
  title={LAGEN: Are Vision-Language Agents Latency Aware?},
  author={Talha Chafekar and Zihan Wang and Zeju Li and Xinyuan Li and Qineng Wang and Jiajun Wu and Yuke Zhu and Yi Dong and Zhiding Yu and Ruohan Zhang and Manling Li},
  year={2026}
}