Language is essential
Removing language or using state-only input increases error substantially, showing that user intent is not a decorative label. It changes the target battery-action trajectory.
EVLA maps energy-field perception, explicit physical state, and natural-language objectives to short-horizon battery-action trajectories.
a Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
b Center for Autonomous Intelligence and e-Mobility, Jeonbuk National University, Jeonju-si, Jeollabuk-do, South Korea
c Dept. of Electrical & Computer Engineering and Computer Science, California State University, Bakersfield, USA
d Hawaii Natural Energy Institute, Honolulu, HI, USA
e Mechanical Engineering, University of Hawaiʻi at Mānoa, Honolulu, HI, USA
f Chicago State University, Chicago, IL 60628, USA
* Corresponding author.
Dataset status: private during review; planned public release after manuscript acceptance.
EVLA adapts the Vision-Language-Action idea to a non-robotic cyber-physical energy setting. Each sample combines an RGB energy-field representation, a physical state vector, a natural-language objective, and a 16-step battery-action reference generated by a physics-consistent shooting-MPC oracle.
The benchmark is built around the controlled relation A = f(S, Z, L), where S is the explicit physical state, Z is a hidden visual-energy regime, L is the language objective, and A is the action trajectory.
Each retained operating state is expanded across three hidden visual-energy regimes and five language objectives, producing 15 counterfactual worlds per base state while preserving the same explicit state.


EVLA does not use ordinary camera scenes. The visual input is an energy-field representation designed to encode hidden operating context through image texture while keeping the explicit state fixed.

The release report checks the integrity of the generated benchmark: malformed JSON files, missing state files, trajectory-length errors, inconsistent SOC or indoor-temperature assignments across counterfactual groups, and sampled missing images.

EVLA is built from processed residential CSV files collected from 19 houses. Each valid 300-sample sliding window is converted into an RGB energy-field representation, paired with an explicit physical state vector, expanded across hidden visual-energy regimes and language objectives, and labeled using a sampling-based shooting-MPC oracle to produce a 16-step battery-action trajectory.

Removing language or using state-only input increases error substantially, showing that user intent is not a decorative label. It changes the target battery-action trajectory.
MobileNet provides a strong accuracy-efficiency tradeoff in the current evaluation, while larger visual backbones do not clearly improve aggregate MSE.
Hidden regimes shift oracle action statistics, but the aggregate learned-model MSE gain from vision is modest in the present release. This is stated directly to avoid overclaiming.



Paper preprint and supplementary material as an ancillary file.
Paper placeholder →Code, scripts, website, small sample data, reproducibility instructions, issues, releases, and changelog.
EVLA-Energy repo →Full benchmark dataset, dataset card, loading examples, and optional model checkpoints.
EVLA dataset →Versioned archival DOI for paper artifacts, dataset metadata, validation logs, and frozen release snapshots.
Coming soonInitial EVLA paper and supplementary package prepared for submission with validated dataset summary, active figures, tables, and ablation bundle.
Hugging Face dataset remains private or restricted; public release is planned after acceptance.
arXiv paper link, supplementary ancillary link, Zenodo DOI, and final public dataset release.
@article{saoud2026evla,
title = {Energy Vision-Language-Action: A Physics-Grounded Multimodal
Benchmark for Intent-Conditioned Residential Energy Management},
author = {Saad Saoud, Lyes and Doukhi, Oualid and Reihani, Ehsan and
Spaci, Saeed and Ghorbani, Reza and Ayyash, Moussa},
journal = {Preprint},
year = {2026},
note = {Initial EVLA dataset release}
}