Distillation¶
Distill any model in the hub into an NN-MTP student. NN-MTP is a compact
machine-learning potential with a ready LAMMPS pair style (nnmtp), so the
student runs large MD on CPUs — no Python and no LibTorch at run time. The
teacher hook is built from the registry's own calculator lines, so any
installed model can be the teacher without teacher-specific code. The
active-learning loop is the separate GPL-2.0 project
onthefly-distill, invoked and
never copied in.
Ask your LLM¶
Distill UMA-s-1p2-OMAT into an NN-MTP student for 300 K NVT MD of slab.vasp.
What you prepare¶
- the teacher — an installed model;
- a structure file with at most four chemical elements;
- a LAMMPS binary with the NN-MTP pair style — the agent can build it once
with
scripts/build_lammps_nnmtp.sh; - an onthefly-distill checkout.
What the engine fixes¶
Know these before you plan a run; changing them means changing the engine, not a setting:
- Student MD — Langevin thermostat at 300 K (NVT-like, time step 0.5 fs). Another temperature or ensemble is not an option of the current engine.
- Boundaries — the input must be fully periodic (a slab keeps its vacuum
along z inside the cell), and the teacher labels are computed periodic. The
student MD and the final LAMMPS check run with
boundary p p fand reflecting walls in z, with the atoms markedFixAtomsheld fixed. The plan states this split; it is the engine's design, not reconciled.
What it asks you¶
- the teacher variant;
- how the student will be used — mainly the MD length you need it to stay stable for; the stability target is set from this;
- a quick trial run or a production run that must meet accuracy and stability targets;
- where LAMMPS and the onthefly-distill checkout are.
What you get¶
- a work directory with the engine config, the teacher hook and
run_distill.sh; - the NN-MTP student, trained where the student failed — teacher labels, student training and student MD in LAMMPS, looped until it survives the target trajectory;
- for a production run, a verdict that needs both held-out accuracy and a stable LAMMPS run within the time budget.
Run it yourself
Run the scripts with the teacher env's interpreter (they import numpy, ase and yaml):
PY=$(python3 -c "import oh_my_mlip; print(oh_my_mlip.resolve('MACE-MPA-0')['python'])")
scripts/build_lammps_nnmtp.sh --repo <onthefly-distill> # once
"$PY" scripts/distill_bootstrap.py --teacher MACE-MPA-0 --structure slab.vasp \
--work ./distill --repo <onthefly-distill> --lmp-bin <lmp>
cd distill && sh run_distill.sh > distill.log 2>&1
--work must be a new or empty directory; nothing in it is overwritten.
For a run you want judged, the bootstrap also takes every target
explicitly (values in < > are yours to choose):
"$PY" scripts/distill_bootstrap.py --teacher MACE-MPA-0 --structure slab.vasp \
--work ./distill --repo <onthefly-distill> --lmp-bin <lmp> \
--acceptance --mode production \
--energy-mae-max <meV/atom> --force-mae-max <meV/A> \
--target-ps <stable MD length, ps> --max-iter <AL rounds> --no-progress-limit <rounds>
and afterwards:
"$PY" scripts/distill_verify.py --work ./distill --json
--mode fixture (at most 3 rounds) only exercises the loop: even a passing
fixture says nothing about the student's accuracy. Use --mode production
for a student you will use. Production seeds the pool with a longer teacher
MD by default (--pool-steps 6000 --pool-save-every 5, about 1200 frames);
the quick demo's 20 frames are too few for the student's energies to meet an
accuracy target.
Full procedure:
recipes/distill.md.