Yani Guan

Ph.D. Candidate, Sautet Group, UCLA  ·  Applied Scientist Intern, SES AI

I build symmetry-aware neural networks for chemistry — E(3)-equivariant models over atoms and trajectories, steerable encoders over molecule images, language models that read 3D directions — along with the multi-scale simulations that feed them and the GPU work that makes them trainable.

profile.jpg
Outside the lab — cameras, Shakespeare, long drives.

Research

How the work fits together Chemistry reaches a model in three forms: papers, structure drawings and 3D atoms from multi-scale simulation. Each goes through an encoder matched to its symmetry: a language model for text, SE(2)-equivariant encoders for drawings, E(3)-equivariant networks for atoms. Irreps-to-text attention and contrastive alignment join them in one latent space with retrieval over all three; post-training turns that into a co-scientist that proposes the next simulation. GPU kernel and precision work sits under every stage. DATA Papers methods, results, claims Structure drawings figures in the literature Multi-scale simulation DFT · MD · kMC · P2D ENCODER · MATCHED SYMMETRY Language model tokens · no geometric group SE(2)-equivariant encoders CN steerable CNN · Ouroboros · VERDICT E(3)-equivariant networks polarizable FF · WignerFlow · world model SMILES → ETKDG → MACE-OFF ALIGNMENT Irreps-to-text YLM · text queries l ≥ 1 Contrastive alignment one latent space Multimodal RAG retrieval over all three POST-TRAINING & AGENT Reasoning traces distilled with GPT-5.5 SFT · LoRA Qwen3.6-27B · GLM-4.7 CO-SCIENTIST reads · plans · runs jobs proposes the next simulation — and submits the jobs that check it COMPUTE GPU Clebsch–Gordan tensor products · channel-wise (uvu) kernels · cuEquivariance · torch.compile · bf16 / tf32 · V100 · A100 · L40S

Scroll the diagram sideways →

Rows are the three forms chemistry reaches a model in; columns are what happens to each. Every form gets an encoder that respects its symmetry — none for text, planar rotations for drawings, E(3) for atoms — before alignment puts them in one space and post-training turns that space into a co-scientist, whose proposals go back to simulation. The GPU work is the floor all of it stands on.

Multi-scale simulation
One ladder from an electrode–electrolyte interface up to a working cell: electronic structure at the surface, molecular dynamics in the liquid, kinetic Monte Carlo for the reaction network, P2D models for the device. Most of the difficulty is in what has to survive the step between rungs.
Electrolyte force fields
Polarizable force-field parameters predicted from molecular structure rather than fitted by hand, which turns weeks of per-chemistry expert work into a forward pass. A formulation goes in as SMILES and comes out with density, conductivity and solubility.
Trajectory world models
Most molecular representations describe a structure; this one describes motion. Latent-space pretraining over MD trajectories, decoded hierarchically from frames to center of mass to atoms, so the learned state carries how an electrolyte moves and not only what is in it.
WignerFlow replaces the MD integrator with an autoregressive E(3)-equivariant transformer: from the last k frames it emits the frame n·δt ahead, directly. Each frame is encoded with eSCN-style SO(2) convolutions — a Wigner-D rotation aligns every edge to the z-axis, which collapses the lmax = 2 tensor product from O(L6) to O(L3) — and causal attention over time draws its logits from invariants only, so the stack stays exactly equivariant. Two heads, MSE regression and equivariant flow matching in displacement space, trained with pushforward unrolling on a curriculum so 104-step rollouts stay on the data manifold; judged on RDF, VACF/VDOS, Li⁺ solvation-shell residence times and energy drift against the reference force field.
Vision–language chemistry
Fine-tuned VLMs read structures straight out of document images (image→SMILES). VERDICT turns several independent recognizers into a consensus engine that votes on molecular identity and abstains when they disagree — on real literature figures, knowing when to refuse is worth more than another point of accuracy.
Ouroboros carries the OCSR work forward with a symmetry prior instead of more data: steerable CN CNNs (escnn, N ∈ {4, 8, 16}) and group-equivariant self-attention as the image encoder, with the SMILES decoder held fixed. Rotations only, never reflections — mirroring a wedge/hash drawing inverts every stereocenter, so a DN-invariant encoder would be blind to chirality. Arms are FLOP-matched (10.3 GFLOPs per image; the parameter-matched C8 would cost 127), equivariance holds to 10−6 relative error on the 90° grid in fp32, and the hypotheses on the synthetic-to-real gap were pre-registered before the first training run. Recognition errors are then pushed through ETKDG and MACE-OFF to price what a wrong diastereomer costs in energy.
Molecular language models — MolT5, 3D-MoLM, even EquiLLM — reduce geometry to invariants before the LLM sees it, so they can name a dipole but not point one. YLM lets text tokens query l ≥ 1 irreps directly: attention logits come from invariants, values carry spherical-harmonic features, and anything with a direction is only scaled or combined through Clebsch–Gordan products. Rotate the molecule by R and the answer rotates by D(R), by construction. Each token runs two streams, an invariant hidden state and an irreps side-stream; parity labels keep (R)- and (S)- apart; a <VEC> / <TENSOR> token hands off to an equivariant readout. Measured on TensorQA, 85k structure–question pairs with GFN2-xTB force and dipole labels; equivariance error below 10−9 in float64.
Scientific multimodality
Text, figures and simulation output are three views of one chemistry that normally sit in three disconnected systems. Contrastive learning with cross-scale alignment puts them in a single latent space: the molecule drawn in a figure, the DFT numbers computed for it, and the way it behaves in an MD box become one object.
Chemistry-aware post-training
Reasoning traces distilled with GPT-5.5 supply the supervision, while chemistry-aware tokenization and a chemistry-aware latent softmax let the model see chemistry where a general tokenizer sees characters. Both feed LoRA SFT on Qwen3.6-27B and GLM-4.7.
Co-scientist
The pieces above compose into a tool-using agent for electrolyte and cell design: it reads the literature, queries the simulation stack, and proposes what to run next. What decides whether it can be trusted is the unglamorous half — leakage-safe benchmarks, and default-deny governed job submission.
Equivariant models on GPUs
Equivariant networks spend their time in Clebsch–Gordan tensor products, so throughput and memory are designed, not discovered. The fully connected e3nn product with per-edge weights was ~40× too slow and gave way to channel-wise (uvu) products; activation memory is budgeted from saved-tensor bytes before a job is queued (the C8 attention encoder fits batch 137 on a 40 GB A100 in bf16); and the profiling stage pits eager PyTorch against torch.compile and cuEquivariance, and fp32 against tf32/bf16 with l ≥ 1 kept in fp32 — on V100s at SDSC Expanse and L40S/A100 nodes on UCLA Hoffman2.

News

Oct 01, 2026 Started three projects on equivariant networks for chemistry: WignerFlow, an autoregressive E(3)-equivariant transformer that generates MD trajectories in place of the integrator; YLM, irreps-to-text attention that lets language tokens query l ≥ 1 geometric features directly; and Ouroboros, steerable SE(2)-equivariant encoders for chemical structure recognition. On the GPU side, channel-wise tensor products replaced e3nn’s fully connected ones for a ~40× speed-up, and the first training runs are queued on V100 nodes at SDSC Expanse and UCLA Hoffman2.
Jun 08, 2026 Excited to join SES AI Corp as an Applied Scientist Intern for the summer!
May 29, 2025 Received Dissertation Year Award at UCLA
Oct 07, 2024 Presented my work on Cu dissolution at PRiME 2024 Conference in Honolulu, Hawaii
Aug 18, 2024 Presented my work on Cu dissolution at the ACS Fall 2024 conference in Denver, CO

All news

Experience

2026
SES AI Applied Scientist Intern
Molecular foundation models, multimodal chemistry models, electrolyte MD, and cell co-scientists for AI-driven battery development.
2022 — present
UCLA, Sautet Group Ph.D. Research
DFT and multi-scale modeling of electrochemical interfaces, electrode degradation, and reaction selectivity; LLM- and vision-based tooling for computational chemistry.
2022
DP Technology Algorithm Researcher Intern
Production ML pipelines for neural network potentials and scalable scientific computing workflows; developer-community and documentation work in the DeePMD ecosystem.
2018 — 2022
Hebei University of Technology Undergraduate Research
Multi-scale simulation (DFT–kMC), machine learning for catalysis, and electrochemical experiments on bifunctional catalysts for Li–S and Zn–air batteries; three first-author publications.

Selected publications

  1. OCSR VLM
    ocsr_real_data.png
    Real Data Closes Synthetic-to-Real Gap in Optical Chemical Structure Recognition
    Yani Guan, Dengpan Dong, Zi Wei, and 4 more authors
    arXiv preprint arXiv:2608.09100, 2026
    * Equal contribution. Vision-language models for image-to-SMILES: real labeled depictions, not more synthetic data, are what close the synthetic-to-real gap.
  1. VERDICT
    VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR
    Yani Guan, Dengpan Dong, Shuang Luo, and 6 more authors
    arXiv preprint arXiv:2608.22183, 2026
    Label-free verification for image-to-SMILES at scale: agreement among four architecturally distinct recognizers reaches AUROC 0.916, while re-rendering a prediction and comparing pixels barely beats chance at 0.547.
  1. MU-PFF MD
    mu_md_platform.png
    High-throughput Molecular Dynamics Simulation on an AI-powered Platform
    Dengpan Dong, Yani Guan, Shuang Luo, and 5 more authors
    ChemRxiv preprint, 2026
    MU-PFF: a machine-learned many-body polarizable force field plus automated system building and trajectory analysis, taking an electrolyte formulation from SMILES to predicted properties without manual parameterization.
  1. Nature Catal.
    Desolvated cations promote CO2 electroreduction through partial covalent interactions
    Yu Shan, Maximilian Jaugstetter, Julian Feijóo, and 14 more authors
    Nature Catalysis, Sep 2026
    Yu Shan, Maximilian Jaugstetter, Julian Feijóo and Yani Guan contributed equally.
  1. Ag Amination
    ag_amination.jpeg
    Electrochemical Amination of Acetone using Ag as Cathode: Mechanism and Role of Pb Impurities for Hydrogen Transfer
    Yani Guan, Justus Kümper, Angelina Cuomo, and 6 more authors
    Catalysis Science & Technology, 2026

All publications