Open-source · Local-first

From failed runs to verified fix

WatcherML is the reliability and recovery layer for ML training. It captures evidence, diagnoses failures, proposes bounded interventions, reruns controlled trials, and proves which change actually worked.

Deterministic capsuleVersioned failure evidence, no LLM required
Bounded executionHard trial, timeout, and GPU budgets
Verifier-owned truthFull trials plus independent confirmations
campaign / llama-lora-oom
recovery active
CAMPAIGN

Recover & verify

validation loss 0.421
SourceCUDA OOM
Trials4 / 6
Confirmations2 required
StatusVerified
Campaign evidence RECORDED
01

Sealed source capsule and contract.

02

Probed batch 16 in a fresh process.

03

Confirmed the candidate twice.

Validation losswithin +0.030
PhaseInterventionResultEvidence
probebatch 32 → 1630 stepssurvived
fullbatch 16 + accum 2completedprovisional
confirmsame sealed config2 / 2 passedverified
↳
Verified recovery2 confirmations passed
Official PyPI install
01 / 04
Core SDK

Install deterministic OOM forensics.

Get the recorder, CLI, failure capsules, bounded recovery policy, and isolated trial runner—without installing the notebook extension or web UI.

local-first SDK + CLI Python 3.10+
terminal
pip install watcherml
Choose your surface

Add notebooks, the local UI, or both.

Keep the core package lean. Add notebook support for Jupyter and Colab, or install the optional local web interface for desktop workflows.

Jupyter + Colab optional local UI CLI always available
optional extras
# Jupyter, IPython, or Colab
pip install "watcherml[notebook]"

# Local web UI
pip install "watcherml[ui]"

# Install both
pip install "watcherml[notebook,ui]"
Python SDK

Wrap the run. Preserve the evidence.

Record configuration, metrics, environment, dataset identity, and runtime telemetry. If training fails, WatcherML freezes the evidence into a deterministic failure capsule.

small API runtime evidence OOM capsule
train.py
import watcherml as watcher

with watcher.init(
    project="mistral-lora",
    config=config,
) as run:
    train(
        model,
        dataset,
        run=run,
    )
Controlled OOM recovery

Rerun, intervene, and verify.

Reproduce the original OOM as a control, evaluate bounded interventions in fresh processes, and require confirmation before any recovery is marked verified.

init runs inspect failures compare export recover recoveries recovery ui
watcher CLI
# Inspect the captured OOM
watcher inspect RUN_ID

# Seal constraints before compute
watcher prepare-recovery RUN_ID \
  --entrypoint train:train \
  --metric validation_loss:minimize:0.42:0.03 \
  --minimum-progress-steps 1000 \
  --confirmation-runs 2 \
  --out recovery-plan.json

# Review, then run bounded trials
watcher recover --plan recovery-plan.json

# Inspect the complete campaign
watcher recovery CAMPAIGN_ID
WHY WATCHERML EXISTS

The crash is one line.
The cost is everything before it.

CUDA Out of Memory (OOM) errors are arguably the most common bottleneck in AI dev. They can erase hours, or even days, of GPU training.

W

WatcherML turns the crash into a bounded recovery protocol. It captures a versioned failure capsule, seals constraints before compute, evaluates typed interventions in fresh processes, and lets independent confirmation—not a promising trial—decide recovery.

See the recovery loop →
MASSIVELY IMPROVE YOUR ML WORKFLOWS

Stop paying full price.
Save your GPU-hours.

WatcherML assists ML engineers when training is expensive, failures recur, multiple engineers are involved, and successful retries must be trusted later. Trackers show you what failed. WatcherML investigates why.

THE BOTTOM LINE

Manual retries can make one ML run pass. WatcherML finds what failed, proposes recovery steps, and executes verified trials.

Designed to work with the stack you already use

PyTorchJupyterNVIDIAMLflowWeights & Biases
V1 PROOF VERTICAL · CUDA OOM RECOVERY

Turn one CUDA OOM into a controlled recovery experiment

WatcherML is a local-first reliability layer for ML training. V1 starts with tackling OOM errors that can be captured, changed, rerun, and verified without guesswork.

watcherml / capture
LOCAL · DETERMINISTIC
01

Capture the evidence that actually existed at failure.

Persist a versioned capsule with the traceback, configuration, last logged step, recent metrics, sampled resources, environment fingerprint, Git state, dataset fingerprint, and optional CUDA allocator bytes.

schema 1.0tracebackallocator statecapture score
1capsule.schema.version = "1.0"
2capsule.failure.class = "cuda_out_of_memory"
3capsule.evidence.training_state.last_logged_step = 417
4capsule.capture = { score: 8, maximum: 10 }
LOCAL-FIRST SDK + CLI

Capture locally. Recover through Python or the CLI.

The recorder works in scripts, Jupyter, and Google Colab. Recovery uses an importable training entrypoint so every probe, full trial, and confirmation can start in a fresh supervised process.

Available through pipNo hosted service, web UI, LLM, or container runtime is required.
import watcherml as watcher

with watcher.init(
    project="mistral-lora",
    config=config,
) as run:
    for step, batch in enumerate(loader):
        loss = train_step(model, batch)
        run.log_metric("loss", float(loss), step=step)

# On failure, the original exception still propagates.
# WatcherML has already persisted the deterministic capsule.
pip install watcherml

# Local terminal—or prefix commands with ! in Colab
watcher init
watcher runs --project mistral-lora
watcher failures --project mistral-lora
watcher inspect RUN_ID
watcher compare FAILED_RUN SUCCESSFUL_RUN
watcher export RUN_ID --out failure-capsule.zip

# Optional local interface; not required in Colab
pip install 'watcherml[ui]'
watcher ui --port 7331
CONTROLLED BY CONSTRUCTION

A recovery trial is evidence, not a guess.

WatcherML v1 cannot rewrite training code, install dependencies, alter datasets, or silently keep searching. It materializes a typed, contract-approved intervention in a fresh process and records exactly what happened.

See the recovery contract →
01Fresh processesEach trial receives a clean interpreter and an independent evidence record.
02Source run preservedThe failed run and its capsule are immutable campaign inputs.
03Allowlisted patchesV1 changes declared configuration keys—not code, packages, or data.
04Hard stop conditionsTrial count, timeout, and wall-clock limits end the campaign predictably.
05Deterministic verdictsProgress, resource, metric, and confirmation checks decide recovery.
BUILT IN PUBLIC

A real MLOps system, tested on Nvidia GPUs.

CaptureVersioned deterministic evidence
BoundTrial, timeout, and GPU budgets
IsolateFresh subprocess per attempt
VerifyIndependent confirmation owns the verdict
Open-source · installable with pip

Help build the experiment engineer your GPU deserves.

Have new ideas and want to contribute? Open a new pull request on Github!