EEVLA
Physics-grounded multimodal benchmark

Energy Vision-Language-Action for residential energy

EVLA maps energy-field perception, explicit physical state, and natural-language objectives to short-horizon battery-action trajectories.

Battery-action trajectory · horizon = 16 steps step 00 · a = +0.00
charge (a > 0) discharge (a < 0) action = 0 baseline
schematic / illustrative — not measured results
A= f ( S · state Z · vision L · language )

Lyes Saad Saouda,*, Oualid Doukhib, Ehsan Reihanic, Saeed Spacid, Reza Ghorbanie, and Moussa Ayyashf

a Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
b Center for Autonomous Intelligence and e-Mobility, Jeonbuk National University, Jeonju-si, Jeollabuk-do, South Korea
c Dept. of Electrical & Computer Engineering and Computer Science, California State University, Bakersfield, USA
d Hawaii Natural Energy Institute, Honolulu, HI, USA
e Mechanical Engineering, University of Hawaiʻi at Mānoa, Honolulu, HI, USA
f Chicago State University, Chicago, IL 60628, USA

* Corresponding author.

Dataset status: private during review; planned public release after manuscript acceptance.

0multimodal samples
0base operating states
3 × 5regimes × objectives
0battery-action steps
0reported ablation runs
01Overview

Energy context and user intent to battery control

EVLA adapts the Vision-Language-Action idea to a non-robotic cyber-physical energy setting. Each sample combines an RGB energy-field representation, a physical state vector, a natural-language objective, and a 16-step battery-action reference generated by a physics-consistent shooting-MPC oracle.

The benchmark is built around the controlled relation A = f(S, Z, L), where S is the explicit physical state, Z is a hidden visual-energy regime, L is the language objective, and A is the action trajectory.

Sample contents

  • RGB energy-field image
  • Physical state: price, SOC, indoor temperature, hour
  • Language objective: energy saving, comfort, balanced, grid support, eco
  • 16-step normalized battery-action trajectory
  • Metadata for controlled counterfactual grouping
02Dataset construction

Balanced counterfactual multi-world expansion

Each retained operating state is expanded across three hidden visual-energy regimes and five language objectives, producing 15 counterfactual worlds per base state while preserving the same explicit state.

EVLA multi-world dataset expansion flow
EVLA expands base operating states into multimodal counterfactual worlds.
Balanced Z by L multi-world expansion matrix
Balanced hidden-regime and language-objective construction. Each cell contains 439,203 samples.
03Energy-field observations

A visual modality for latent energy context

EVLA does not use ordinary camera scenes. The visual input is an energy-field representation designed to encode hidden operating context through image texture while keeping the explicit state fixed.

Representative EVLA energy-field samples
Representative energy-field samples across hidden regimes and language objectives.
04Validation

Structural checks for reproducibility

The release report checks the integrity of the generated benchmark: malformed JSON files, missing state files, trajectory-length errors, inconsistent SOC or indoor-temperature assignments across counterfactual groups, and sampled missing images.

✓0
bad JSON files
✓0
missing state files
✓0
trajectory-length errors
✓0
inconsistent SOC/Tin groups
✓0 / 6,588
missing sampled images
Physical-state distributions in EVLA
Physical-state distributions for SOC, indoor temperature, price, and hour.
05Dataset generation

Physics-grounded EVLA construction pipeline

EVLA is built from processed residential CSV files collected from 19 houses. Each valid 300-sample sliding window is converted into an RGB energy-field representation, paired with an explicit physical state vector, expanded across hidden visual-energy regimes and language objectives, and labeled using a sampling-based shooting-MPC oracle to produce a 16-step battery-action trajectory.

EVLA dataset generation pipeline
EVLA dataset generation pipeline. Processed residential time-series windows are transformed into multimodal tuples containing the energy-field image V, physical state S, language objective L, hidden visual-energy regime Z, and oracle battery-action trajectory A.
06Results

Main empirical findings

FINDING / 01

Language is essential

Removing language or using state-only input increases error substantially, showing that user intent is not a decorative label. It changes the target battery-action trajectory.

FINDING / 02

Compact backbones compete

MobileNet provides a strong accuracy-efficiency tradeoff in the current evaluation, while larger visual backbones do not clearly improve aggregate MSE.

FINDING / 03

Vision encodes hidden context

Hidden regimes shift oracle action statistics, but the aggregate learned-model MSE gain from vision is modest in the present release. This is stated directly to avoid overclaiming.

Mean oracle trajectories by objective
Language-conditioned oracle action profiles.
Backbone-selection MSE
Backbone-selection MSE over three random seeds.
MobileNet modality ablation
MobileNet modality ablation.
07Release plan

Where each EVLA artifact lives

GitHub

Code, scripts, website, small sample data, reproducibility instructions, issues, releases, and changelog.

EVLA-Energy repo →

Hugging Face

Full benchmark dataset, dataset card, loading examples, and optional model checkpoints.

EVLA dataset →

Zenodo

Versioned archival DOI for paper artifacts, dataset metadata, validation logs, and frozen release snapshots.

Coming soon
—News

Initial EVLA paper and supplementary package prepared for submission with validated dataset summary, active figures, tables, and ablation bundle.

Hugging Face dataset remains private or restricted; public release is planned after acceptance.

arXiv paper link, supplementary ancillary link, Zenodo DOI, and final public dataset release.

08Citation

BibTeX

@article{saoud2026evla,
  title   = {Energy Vision-Language-Action: A Physics-Grounded Multimodal
             Benchmark for Intent-Conditioned Residential Energy Management},
  author  = {Saad Saoud, Lyes and Doukhi, Oualid and Reihani, Ehsan and
             Spaci, Saeed and Ghorbani, Reza and Ayyash, Moussa},
  journal = {Preprint},
  year    = {2026},
  note    = {Initial EVLA dataset release}
}