The AI model that can see science.

Ask an LLM where to drill the next oil well or which drug binds a protein, and it cannot answer: it cannot see the field or the molecule. It works through tools and their text output, so an expert still has to look at the data and decide the next step. SciLM, a multimodal language model, sees the data itself, makes the first guess the way an expert would, and drives the loop autonomously with fewer expensive simulations or experiments.

Today we generate physics-based simulation data to order and train specialized AI models. SciLM comes next.

107 mEarth

LLMs read digits, SciLM sees physics.

An LLM receives a molecule as a string of numbers and letters, or a reservoir as a grid of numbers, and works through tools whose answers come back as text. SciLM, the Scientific Language Model, is trained from scratch to read and write scientific objects in their native form, so the tools live inside it and it answers scientific questions no one wrote a tool for.

tool calltool calltool calltool callProtein foldingFlow simulatorMolecular dynamicsa tool it writesitself, untestedLLMreasons and plans withwordssimulates on its ownnothing
An LLM with tools. Every answer comes back as text, and reliably only for questions a tool already exists for.
can still callany tool,if neededSciLMreasons and plans withwordsscientific objectssimulates on its own, the tools in its weightsProtein foldingFlow simulatorMolecular dynamicsanswersquestions no tool exists for
SciLM. The tools live in its weights: protein folding, a flow simulator, molecular dynamics, in one model. It plans with the science in mind, and answers questions no tool exists for. It can still call the same tools if needed, but it can also read their results in their native form and not only text summaries.

Read the thesis

AI needs more data than science has ever measured. We generate it.

The breakthroughs in AI were built on large datasets that already existed: the web for ChatGPT, the Protein Data Bank for AlphaFold. Measurements from different labs rarely combine into one training set, and much of what matters cannot be measured at all: no experiment watches a reaction at the atomic scale or sees flow in the subsurface. Computational science has generated what experiments cannot see for seventy years, and it can do so at larger scales and in greater quantities.

The webLLMs

Trillions

of words

Written by people, for people.

Protein Data BankAlphaFold

240,000

structures in 55 years

Experimentally measured, one at a time.

SimulationSciLM

102.4 M

structures, generated

More than every laboratory in history has measured experimentally. Plus 1,000,000 oil and gas reservoirs.

SciLMshape of the physicsSimulationpretrain: every scale, any quantityExperimentfine-tune: correct the small errorsPredictionsfor questions no tool exists for
Pretrain on simulation. Fine-tune on experiment. Nobody can see a reaction at the atomic scale, and nobody can see the subsurface, so scientists and engineers simulate them. An AI model has to learn the same way to gain a deep understanding of the physical world.

Released so far.

MaterialsSaddles

Reactant, transition state and product for reactions across inorganic materials.

34.14 Mreactions
102.4 Mstructures
687 GBCC-BY 4.0

SiliciclasticReservoirs

Synthetic 3D oil and gas reservoirs across eight depositional architectures, with facies, porosity and permeability per voxel.

1,000,000reservoirs
8architectures
787 GBCC-BY 4.0

The engines, the models, and data to order

Data first, models now, SciLM next. Every step pays.

One model for all of science needs data no one else has. So SciLM generates the data first, builds specialized AI models on it next, and trains SciLM on top. Each step is a product on its own: datasets and AI models to license, and data generated to order for your domain ranging from kilometers of subsurface rock down to atomic scale. No one else trains a model on both kinds of data to reason over both text and scientific objects.

First

Data

Generated by physics-based simulation, with the domain expertise of geoscientists, physicists, chemists and engineers built in. The first datasets are open; the next will be licensed.

34.14 M reactions1,000,000 subsurface modelsSaddleMill, ResMill, PotMill

Now

Specialized AI models

One model per data type, each trained to optimize prediction and generation for its own physical system. Licensed as tools to LLMs or on their own for research and engineering design.

ResFlow, oil and gas reservoirsSaddleFlow, inorganic materialsthe next domains

Next

SciLM

True generalization with one model: it reads and writes every form of data needed for science and engineering, and reasons over text and scientific objects in one shared representation.

textatomsvolumesfields

Founded by two computational scientists/engineers.

Ilgar Baghishov

Ilgar Baghishov

Co-founder and CEO

Wrote a PhD dissertation at UT Austin on generating the training data AI in science needs. Created SaddleMill and ResMill, the engines behind both SciLM datasets, and SiliciclasticReservoirs, one million 3D reservoir models. Trained SaddleFlow, the first generative model for transition states in inorganic materials, and ResFlow, the first single model for every siliciclastic reservoir type, any wells, any field size. Built PotMill at Los Alamos National Lab, which automates ML interatomic potentials from data generation to the final model in under one day. Petroleum engineer by training.

Sung Hoon Jung

Sung Hoon Jung

Co-founder

Built agentic systems for computational scientists and engineers as part of his PhD at UT Austin. Generated MaterialsSaddles, 34M transition states on 500k GPU-hours. Refactored the UF3 ML interatomic potential to cut peak training memory 223x and model size 62x with no accuracy loss. Discovered a gel electrolyte cycling zinc batteries past 3,600 hours. Architected and administers three HPC clusters totaling 150 nodes. Biomedical engineer by training, with US and European patents pending on an endotracheal tube securement device.

Scientific advisors

Graeme Henkelman

Graeme Henkelman

Professor of Chemistry, Oden Institute for Computational Engineering and Sciences, UT Austin

Developed the nudged elastic band and dimer methods and the Bader charge analysis code: the standard ways to find transition states and charges in materials simulation. Cited more than 100,000 times.

Michael Pyrcz

Michael Pyrcz

Professor of Petroleum and Geosystems Engineering, Jackson School of Geosciences and Oden Institute for Computational Engineering and Sciences, UT Austin

Spent thirteen years with Chevron’s Energy Technology Company before joining UT Austin, modeling the subsurface and conducting and leading research in energy data science and AI. Authored the textbook Geostatistical Reservoir Modeling (Oxford), the e-books Applied Geostatistics in Python and Applied Machine Learning in Python, and created the GeostatsPy Python package. As the GeostatsGuy on YouTube, teaches geostatistics and machine learning to tens of thousands of learners worldwide.

John T. Foster

John T. Foster

Professor of Petroleum and Geosystems Engineering, Oden Institute for Computational Engineering and Sciences, UT Austin

Spent seven years at Sandia National Laboratories as a core developer of Peridigm, the massively parallel code for simulating how materials fracture. Wrote the open course books on numerical methods and high-performance computing his UT Austin students learn from, and co-founded daytum.

If the data or model you need does not exist yet, write to us.

We generate simulation data and train AI models to specification: the physics, the systems, the sampling and the format, for the subsurface and for materials alike. Exclusive or non-exclusive licenses.

  • Datasets: MaterialsSaddles and SiliciclasticReservoirs are open; the next are licensed.
  • Models: ResFlow and SaddleFlow, licensed as tools or trained on your data.
  • Data to order: your system, your scale, in the form your models read.