PhD Student · Machine Learning · University of Pennsylvania

Wenrui Ma

Learning representations of dynamical-system data that generalize across subjects, devices, and sensor configurations.

Portrait of Wenrui Ma
About

Hi, I’m Wenrui

I’m a first-year PhD student in machine learning, advised by Dr. Eva Dyer in the NerDS Lab. I build models that learn transferable structure from messy real-world time series, from neural recordings to wearable sensors.

Before my PhD, I earned an M.S. in Mathematics and an M.S. in Computer Science at Georgia Tech, a B.S. in Computer Science at Colorado State University (Magna Cum Laude), and a B.Eng. in Software Engineering at Hunan University.

Research

A unified view of time series

Foundation models transfer easily across images and text, but time series (the native language of biology and embodied systems) still resist it. I treat neural recordings, wearable sensor streams, robot trajectories, and macro-scale series as different observations of underlying dynamical systems, each measured through some configuration of channels.

My work builds representations that disentangle a system's temporal dynamics from its measurement-channel geometry, so one pretrained model can adapt to new subjects, devices, and channel sets, and serve both classification and forecasting. I work across biological signals (EEG and intracranial recordings, currently cross-subject iEEG in OCD), physiological and robot-trajectory data, and macro-scale financial series.

Publications

Papers

Figure from NoTS
ICLR 2026 · TSALM Workshop · Oral
Narrative of Time across Scales (NoTS)
Wenrui Ma, Ran Liu, Ellen L. Zippi, Christopher M. Sandino, Juri Minxha, Behrooz Mahasseni, Erdrin Azemi, Ali Moin, Eva L. Dyer
Reframes time-series pretraining as autoregressively predicting the next scale in a coarse-to-fine decomposition of the signal — a narrative across resolutions rather than across time steps. Lightweight channel and task prompts then adapt one frozen backbone to unseen sensors and to tasks like classification, imputation, and anomaly detection.
Figure from Multi-Level Transport Alignment
Preprint
Multi-Level Transport Alignment for Time-Series Domain Generalization
Wenrui Ma, Zihao Chen, Ran Liu, Sergey A. Shuvaev, Eva L. Dyer
Most domain-generalization methods align time series pair-by-pair; MULTA instead learns a single balanced optimal-transport plan over the whole batch, organizing class and domain structure globally. An alignment-gap reweighting concentrates training on the samples whose learned structure strays furthest from that target, improving cross-subject and cross-dataset transfer.
Figure from BrainWideBench
Preprint
BrainWideBench: A Large-Scale, Multi-Task Benchmark for Neuro-Foundation Models
Alexandre Andre, S. P. Mahato, Vinam Arora, Wenrui Ma, et al.
A common yardstick for brain-wide foundation models: one model pretrains across 126 mice and 423 sessions spanning behavior, spiking activity, and anatomy, then transfers to held-out animals on a battery of decoding tasks. It standardizes how “general-purpose” neural models are trained and compared.
Figure from SCRYER
Preprint
SCRYER: A Scalable Framework for Forecasting Neural Population Activity
Zihao Chen, Stephen Kwak, Alexandre Andre, Ian J. Knight, Wenrui Ma, Bijan Pesaran, Sergey A. Shuvaev, Eva L. Dyer
Neural activity is dominated by slow background drift, so forecasters trained on average error tend to miss the sparse spikes that actually carry information. SCRYER’s event-driven contrastive objective targets those transients, while unified neuron embeddings and cross-neuron attention handle populations that change between sessions — producing rare positive transfer across species.
Figure from Mixtures of Localized Diffusion Processes
Preprint
Mixtures of Localized Diffusion Processes for Time Series Forecasting
Zihao Chen, Alexandre Andre, Wenrui Ma, Ian J. Knight, Mehdi Azabou, Anqi Wu, Sergey A. Shuvaev, Eva L. Dyer
Running diffusion over an entire forecast is wasteful when most of a signal is easy to predict. MeLD decomposes each signal into learnable frequency bands, scores how uncertain each band is, and applies diffusion only to the uncertain bands while confident bands pass straight through an MLP.
Figure from PRISM
Preprint
PRISM: A Hierarchical Multiscale Approach for Time Series Forecasting
Zihao Chen, Alexandre Andre, Wenrui Ma, Ian J. Knight, Sergey A. Shuvaev, Eva L. Dyer
Splits each signal hierarchically in time and frequency into a bank of bands, then learns selection weights that decide which resolutions matter before recombining them. The result is one forecaster that captures slow trends and fine detail together, instead of committing to a single fixed scale.
Figure from the intracranial recordings paper
NeurIPS 2025 · BrainBodyFM Workshop
A Scalable Self-Supervised Method for Modeling Human Intracranial Recordings During Natural Behavior
S. P. Mahato, Jingyun Xiao, Alexandre Andre, Geeling Chau, Wenrui Ma, Ian J. Knight, et al., Eva L. Dyer
Intracranial datasets are hard to pool because electrode layouts differ from person to person. A Perceiver-based model masks a subset of channels and reconstructs them from the rest via learnable embeddings, learning label-free representations that scale with participants and improve decoding of behavior, speech, and audition.
Figure from Balanced Data, Imbalanced Spectra
ICML 2024
Balanced Data, Imbalanced Spectra: Unveiling Class Disparities with Spectral Imbalance
Chiraag Kaushik, Ran Liu, Chi-Heng Lin, Amrit Khera, Matthew Y. Jin, Wenrui Ma, Vidya Muthukumar, Eva L. Dyer
Even when every class has equally many training examples, their learned feature covariances have different eigenspectra — a “spectral imbalance” that sample counts can’t reveal. The paper shows this geometric disparity, not class frequency, is what tracks the accuracy gaps between classes.
Contact

Get in touch

I’m always glad to connect. Whether you’d like to discuss ideas in time-series representation learning and biosignals, explore a collaboration, or just ask a question, feel free to reach out.