Rows are the three forms chemistry reaches a model in; columns are what happens to each. Every form gets an encoder that respects its symmetry — none for text, planar rotations for drawings, E(3) for atoms — before alignment puts them in one space and post-training turns that space into a co-scientist, whose proposals go back to simulation. The GPU work is the floor all of it stands on.
Multi-scale simulation
One ladder from an electrode–electrolyte interface up to a working cell: electronic structure at the surface, molecular dynamics in the liquid, kinetic Monte Carlo for the reaction network, P2D models for the device. Most of the difficulty is in what has to survive the step between rungs.
Electrolyte force fields
Polarizable force-field parameters predicted from molecular structure rather than fitted by hand, which turns weeks of per-chemistry expert work into a forward pass. A formulation goes in as SMILES and comes out with density, conductivity and solubility.
Trajectory world models
Most molecular representations describe a structure; this one describes motion. Latent-space pretraining over MD trajectories, decoded hierarchically from frames to center of mass to atoms, so the learned state carries how an electrolyte moves and not only what is in it.
WignerFlow replaces the MD integrator with an autoregressive E(3)-equivariant transformer: from the last k frames it emits the frame n·δt ahead, directly. Each frame is encoded with eSCN-style SO(2) convolutions — a Wigner-D rotation aligns every edge to the z-axis, which collapses the lmax = 2 tensor product from O(L6) to O(L3) — and causal attention over time draws its logits from invariants only, so the stack stays exactly equivariant. Two heads, MSE regression and equivariant flow matching in displacement space, trained with pushforward unrolling on a curriculum so 104-step rollouts stay on the data manifold; judged on RDF, VACF/VDOS, Li⁺ solvation-shell residence times and energy drift against the reference force field.
Vision–language chemistry
Fine-tuned VLMs read structures straight out of document images (image→SMILES). VERDICT turns several independent recognizers into a consensus engine that votes on molecular identity and abstains when they disagree — on real literature figures, knowing when to refuse is worth more than another point of accuracy.
Ouroboros carries the OCSR work forward with a symmetry prior instead of more data: steerable CN CNNs (escnn, N ∈ {4, 8, 16}) and group-equivariant self-attention as the image encoder, with the SMILES decoder held fixed. Rotations only, never reflections — mirroring a wedge/hash drawing inverts every stereocenter, so a DN-invariant encoder would be blind to chirality. Arms are FLOP-matched (10.3 GFLOPs per image; the parameter-matched C8 would cost 127), equivariance holds to 10−6 relative error on the 90° grid in fp32, and the hypotheses on the synthetic-to-real gap were pre-registered before the first training run. Recognition errors are then pushed through ETKDG and MACE-OFF to price what a wrong diastereomer costs in energy.
Molecular language models — MolT5, 3D-MoLM, even EquiLLM — reduce geometry to invariants before the LLM sees it, so they can name a dipole but not point one. YLM lets text tokens query l ≥ 1 irreps directly: attention logits come from invariants, values carry spherical-harmonic features, and anything with a direction is only scaled or combined through Clebsch–Gordan products. Rotate the molecule by R and the answer rotates by D(R), by construction. Each token runs two streams, an invariant hidden state and an irreps side-stream; parity labels keep (R)- and (S)- apart; a <VEC> / <TENSOR> token hands off to an equivariant readout. Measured on TensorQA, 85k structure–question pairs with GFN2-xTB force and dipole labels; equivariance error below 10−9 in float64.
Scientific multimodality
Text, figures and simulation output are three views of one chemistry that normally sit in three disconnected systems. Contrastive learning with cross-scale alignment puts them in a single latent space: the molecule drawn in a figure, the DFT numbers computed for it, and the way it behaves in an MD box become one object.
Chemistry-aware post-training
Reasoning traces distilled with GPT-5.5 supply the supervision, while chemistry-aware tokenization and a chemistry-aware latent softmax let the model see chemistry where a general tokenizer sees characters. Both feed LoRA SFT on Qwen3.6-27B and GLM-4.7.
Co-scientist
The pieces above compose into a tool-using agent for electrolyte and cell design: it reads the literature, queries the simulation stack, and proposes what to run next. What decides whether it can be trusted is the unglamorous half — leakage-safe benchmarks, and default-deny governed job submission.
Equivariant models on GPUs
Equivariant networks spend their time in Clebsch–Gordan tensor products, so throughput and memory are designed, not discovered. The fully connected e3nn product with per-edge weights was ~40× too slow and gave way to channel-wise (uvu) products; activation memory is budgeted from saved-tensor bytes before a job is queued (the C8 attention encoder fits batch 137 on a 40 GB A100 in bf16); and the profiling stage pits eager PyTorch against torch.compile and cuEquivariance, and fp32 against tf32/bf16 with l ≥ 1 kept in fp32 — on V100s at SDSC Expanse and L40S/A100 nodes on UCLA Hoffman2.