Researchers have developed a new AI model called NeuroWorld that can predict how the human brain reacts to movies. Unlike previous models that just guess a single response, this model simulates the brain's internal state over time. This allows it to "roll out" and forecast brain activity into the future even as the movie continues.
The failure of retrospective regression
Predicting how the brain responds to continuous sensory input—like watching a film—is a central challenge in computational neuroscience. Traditionally, this has been approached as a stimulus-to-response regression problem. In this paradigm, researchers attempt to map specific sensory features directly onto the blood-oxygen-level-dependent (BOLD) signal. This signal is the proxy measurement used in functional Magnetic Resonance Imaging (fMRI) to trace neural activity.
However, the authors of this study argue that this approach is fundamentally flawed for simulating real-world experiences. Most existing models, such as TRIBE or MIRAGE, suffer from "future-stimulus leakage." This occurs when the model uses information from future frames to predict current brain activity. This violates the principle of causality. A biological brain can only react to what it has already perceived, not what is coming next .
Furthermore, these models often focus on reconstructing the noisy fMRI signal itself. They do not prioritize learning the underlying "latent" state. A latent state is a mathematical abstraction representing the hidden neural dynamics that drive brain evolution. Without capturing this state, a model cannot perform a "rollout." A rollout is the process of recursively feeding its own predictions back into itself to simulate a continuous trajectory.
A two-stage latent world model
To solve this, the researchers propose NeuroWorld. This framework treats brain activity as a stimulus-conditioned evolution within a learned latent space. Instead of trying to rebuild the messy fMRI image, NeuroWorld focuses on the mathematical abstraction of the brain's state. The architecture operates in two distinct stages .
In the first stage, Latent Dynamics Learning (LDL), the model learns to represent the brain's endogenous (internal) state and its causal evolution. An fMRI encoder maps the observed brain activity into a low-dimensional latent vector. Meanwhile, an action encoder converts multimodal stimuli—video, audio, and text—into "stimulus-action tokens." The model is trained via next-latent prediction. It looks at the current latent state and the current stimulus to predict the next latent state.
Crucially, the authors optimize this using a "teacher forcing" method during training. During training, the model uses actual measured data to guide its next step. However, they do not include a reconstruction loss. This means the model is not penalized for failing to recreate the exact pixel-level noise of an fMRI scan. It is only penalized if it fails to predict the correct future state. To prevent the latent space from collapsing, they use Sketched Isotropic Gaussian Regularization (SIGReg). This ensures the representations remain diverse and informative.
The second stage, Latent Rollout Decoding (LRD), turns these abstract mathematical states back into something readable. Once the dynamics are learned, the LDL components are frozen. The model is then given a short "prefix" of actual observed fMRI data to initialize its state. From there, it enters an autoregressive rollout. It repeatedly uses its own predicted latent states and the incoming movie stimuli to "dream" the future trajectory of the brain. Finally, a decoder maps these predicted latents back into whole-brain fMRI responses.
Robustness across diverse benchmarks
The authors evaluate NeuroWorld across three naturalistic movie-fMRI benchmarks involving 30 total participants. To strengthen their empirical foundation, they introduce the Singapore Multimodal Imaging & Naturalistic Dataset (SG-MIND). This is a new benchmark containing 20 participants and over 140 hours of audiovisual viewing .
The results show that NeuroWorld achieves state-of-the-art performance in causal multi-step rollouts. On the SG-MIND benchmark, the paper reports a global Pearson correlation ($r$) of 0.2190. This $r$ value measures how closely the predicted trajectory matches the actual one. It also reports an identification accuracy (Cls@Top10) of 0.8044. This means the model correctly identifies the specific brain trajectory segment in 80.44% of cases. This is a significant jump over regression-based baselines. Those baselines tended to collapse toward chance when forced to operate under strictly causal constraints.
One striking finding is the model's resilience to "autoregressive drift." This is the tendency for errors to compound when a model feeds its own predictions back into itself. Even as the forecasting horizon increases from 20 to 100 time points (TRs), the decline in correlation is relatively gradual [Table III]. This suggests the learned dynamics are stable enough for meaningful long-term simulation. Additionally, the model demonstrates that multimodal conditioning is essential. Combining video, audio, and text provides superior predictive power compared to any single modality .
Limitations in scope and resolution
Several technical trade-offs remain. First, the model's predictive power is not uniform across the brain. The authors find that prediction accuracy is highest in the posterior regions. These include the occipital and temporal cortices, which are heavily driven by visual and auditory stimuli .
In contrast, activity in the anterior association regions remains much harder to forecast. This suggests that while NeuroWorld captures sensory-driven dynamics well, it may not yet fully capture complex, internally-generated thoughts.
Second, the study relies on ROI-level (Region of Interest) signals. Instead of analyzing every individual voxel (a 3D pixel in a brain scan), the researchers partition the brain into roughly 1,000 functional parcels. This makes the computation tractable. However, it sacrifices the fine-grained spatial resolution needed for studying highly localized neural phenomena. Finally, the model currently requires subject-specific "heads" in the decoding stage. These are specialized layers used to map the universal latent dynamics to an individual's unique brain anatomy.
The verdict
NeuroWorld represents a successful pivot from simple pattern matching to true dynamical modeling. By prioritizing the stability of the latent transition over the perfection of the fMRI reconstruction, the authors have built a tool that can simulate the passage of time in the human brain.
For practitioners in computational neuroscience, the takeaway is clear. If you want to build a simulator, stop optimizing for reconstruction accuracy. Instead, start optimizing for transition stability. NeuroWorld is not yet a plug-and-play replacement for all brain encoding tasks. However, it establishes a principled blueprint for moving from "what does this stimulus look like in the brain?" to "how does this stimulus change the brain's trajectory?" Code is reportedly available; see the paper for the canonical link.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 94% (passed)
Claims verified: 20 / 20
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 85,141
Wall-time: 230.3s
Tokens/s: 369.7