Anthropic Interview: Design an ML Configuration System
A full solution to the Anthropic ML configuration question: composition and overrides, typed validation, the resolved config, run records for reproducibility, a working Python loader, and the tradeoffs interviewers probe.
By The Forward Deployed editorial teamReviewed
Part of the Anthropic system design question bank. The question is representative of the round. The analysis and solution are this site's own.
Problem statement
This is a research-track option. It mixes design discussion with a small implementation. The system takes config files and command-line overrides, produces one fully resolved and validated configuration, hands it to the trainer, and records everything needed to relaunch the run later.
Clarifying questions
- How many runs, and how big? Hundreds of runs a week, from one GPU to thousands.
- Who writes configs? Researchers, often in a hurry. Mistakes are common and expensive at scale.
- What must reproducible mean? Same settings, code, data, and environment. Ask whether bitwise-identical results are required. They usually cannot be promised on GPUs.
- Where do runs launch? A cluster scheduler that takes a container image and a command.
- Is there an existing format? Assume YAML files in the training repository.
- Should the round include code? Assume yes: write the loader.
What makes configuration hard
Configs look like a text-file problem. Three things make them a systems problem.
Mistakes are expensive. A typo in a learning rate key silently falls back to a default, and the run burns a thousand GPUs for a day before anyone notices the loss curve looks wrong.
Configs change underneath old runs. Someone changes a shared default next month, and last month's config file now means something different. "Rerun the config file" no longer reproduces the run.
Researchers need speed. A system so strict that every experiment needs a schema change will be bypassed with copy-pasted scripts.
So the driving tension is flexibility versus safety. The design must make the common edit fast and the dangerous mistake impossible.
flowchart LR
F[Config pieces<br/>model, optim, data, cluster]:::svc --> M[Merge in fixed order]:::svc
O[CLI overrides]:::user --> M
M --> V{Validate:<br/>types, unknown keys,<br/>cross-field rules}:::svc
V -->|fail| E[Error before launch]:::bad
V -->|ok| R[Resolved config]:::store --> RR[(Run record)]:::store
R --> T[Trainer]:::svc
classDef user fill:#e6efec,stroke:#315e55,color:#171717;
classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717;
classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717;
classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717;Key idea. Validate before launch, and save the resolved result, never the recipe. The recipe changes; the resolved config does not.
Key concepts
Composition
A run config is assembled from reusable pieces grouped by kind: a model preset such as model/7b, an optimizer preset, a data mix, a cluster layout. The run file names the pieces and adds its own values.
Merge order
Pieces merge in a fixed order, and later values win: defaults, then each piece in the listed order, then the run file, then command-line overrides. Dictionaries merge key by key. Lists replace whole, because merging lists item by item produces surprises nobody can predict.
Overrides
A dotted path sets any value from the command line: optim.lr=3e-4. Values parse as YAML, so true becomes a boolean and 2048 an integer. One trap: YAML 1.1 parsers, including PyYAML, read 3e-4 as a string because it has no decimal point. The schema must coerce and check types, or that string reaches the trainer. Unknown paths are errors, unless the user explicitly adds a new key with a marker such as +new.key=1.
The resolved config
After merging and interpolation, the resolved config has every value explicit and no references left. It is what the trainer receives and what gets saved. Its hash is the run's config fingerprint.
Reproducibility
Reproducing a run needs five things: the resolved config, the exact code commit, the container image digest, the data snapshot, and every random seed. Missing any one of them means the rerun may differ.
Key idea. Compose from pieces, merge in a fixed order, override by dotted path, validate strictly, and freeze the resolved result with the code, image, data, and seeds.
- Requirements
Before reading on. List the requirements. Which property matters most, and what constraint shapes the design?
1.1 Functional requirements
- Compose a run from reusable config pieces.
- Override any value from the command line.
- Validate types, reject unknown keys, and enforce cross-field rules before launch.
- Produce one resolved config and a fingerprint.
- Record everything needed to relaunch the run.
- Relaunch any past run from its record.
- Show the difference between two runs' configs.
1.2 Non-functional requirements
- Fast feedback. Validation runs in seconds on a laptop, before any job reaches the cluster.
- Readable. Configs are plain data that diff cleanly in code review.
- Stable. A change to a shared default cannot change a past run's record.
1.3 The constraint versus the property
Reproducibility is the property. An experiment that cannot be rerun cannot be trusted, compared, or debugged. Researcher speed is the constraint. Every safety feature must cost the researcher almost nothing, or they will route around it.
- Back-of-the-envelope estimation
This is not a scale problem, but two numbers justify the design.
Cost of a silent mistake. For practice: a run on 1,024 GPUs for one day is 24,576 GPU-hours. A typo that falls back to a default wastes all of it. Rejecting unknown keys before launch costs a few milliseconds. That ratio is the argument for strict validation.
Size of the records. A resolved config is a few kilobytes. Hundreds of runs a week produce a few megabytes a year of records. Storing every resolved config forever costs nothing.
- API design
3.1 Command line
train run experiment=exp/scaling_7b optim.lr=3e-4 data.mix=web_v3 train validate experiment=exp/scaling_7b # resolve + validate, no launch train show experiment=exp/scaling_7b # print the resolved config train diff <run_id_a> <run_id_b> # key-by-key difference train relaunch <run_id> [--allow-drift] # rerun from the record
3.2 Library
build(files: list[path], overrides: list[str]) -> Run # merged, validated fingerprint(run: Run) -> str # hash of resolved config record(run: Run) -> RunRecord # adds commit, image, data, seeds load_record(run_id) -> RunRecord
- Data model
4.1 Config pieces on disk
configs/ defaults.yaml model/1b.yaml model/7b.yaml model/70b.yaml optim/adamw.yaml optim/adafactor.yaml data/web_v3.yaml data/code_mix.yaml cluster/a_1024.yaml exp/scaling_7b.yaml # lists pieces + run-specific values
4.2 An experiment file
compose: [defaults, model/7b, optim/adamw, data/web_v3, cluster/a_1024] total_steps: 200000 global_batch: 2048 seed: 17 optim: lr: 3.0e-4 warmup_steps: 2000
4.3 The run record
run_record
run_id, created_at, created_by,
resolved_config (full JSON), config_fingerprint,
git_commit, git_dirty: false,
image_digest,
data_snapshot_id, data_manifest_hash,
seeds: {python, numpy, torch, data_order},
hardware: {gpu_type, gpu_count, parallel_layout},
library_versions: {...}Relaunch reads this record. It never re-reads the config files.
- High-level design
5.1 One Python file of constants
LR = 3e-4 BATCH = 2048 # edit and run
It is fast and it fails every requirement: no composition, no validation, no record. The next person edits the file and the old run is gone.
5.2 Fix 1: YAML pieces and a merge
Move settings into YAML pieces and merge them in a fixed order. Add dotted-path overrides from the command line. Composition now works, and configs diff cleanly in review.
A typo still passes silently: learning_rte: 1e-3 merges fine, and the trainer reads the default lr.
5.3 Fix 2: typed schemas and strict validation
Define the config as typed classes. Building them from the merged dictionary fails on unknown keys and wrong types. Cross-field rules live on the classes and run at build time.
flowchart LR D[Merged dict]:::svc --> S[Typed schema]:::new S -->|unknown key| X1[Error: learning_rte is not a field]:::bad S -->|wrong type| X2[Error: lr must be float]:::bad S -->|rule broken| X3[Error: batch must divide by workers]:::bad S -->|ok| R[Run object]:::store classDef svc fill:#f4f1e8,stroke:#315e55,color:#171717; classDef store fill:#fdf3dc,stroke:#c4492d,color:#171717; classDef bad fill:#fbe9e4,stroke:#c4492d,color:#171717; classDef new fill:#ffffff,stroke:#c4492d,stroke-width:2px,stroke-dasharray:5 3,color:#171717;
5.4 Fix 3: resolve, fingerprint, and record
Serialize the validated object to its resolved form, hash it, and write a run record with the commit, image digest, data snapshot, and seeds before the first training step. The launcher refuses to start from a dirty git tree unless the user passes an explicit flag, which is also recorded.
5.5 The composed design
sequenceDiagram autonumber actor R as Researcher participant CLI as train CLI participant L as Loader participant RS as Run store participant S as Cluster scheduler R->>CLI: train run experiment=... optim.lr=3e-4 CLI->>L: build(files, overrides) L->>L: merge in order, apply overrides L->>L: validate types, keys, rules L-->>CLI: Run (or errors, in seconds) CLI->>RS: write run record (resolved config, commit, image, data, seeds) CLI->>S: launch image with --run-id S->>S: trainer loads the record by run id
The trainer loads the record by run ID, never the config files. What was recorded is exactly what runs.
Key idea. Merge for convenience, validate for safety, and make the saved record the only thing the trainer reads.
- Deep dives
6.1 The loader, in code
Before reading on. Write a loader that merges YAML files, applies dotted overrides, rejects unknown keys, and checks that the global batch divides evenly across workers.
A compact version using dataclasses. It runs as written with PyYAML installed.
from dataclasses import dataclass, field, asdict, fields
import hashlib, json
import yaml
@dataclass(frozen=True)
class Optim:
name: str = 'adamw'
lr: float = 3e-4
warmup_steps: int = 2000
weight_decay: float = 0.1
@dataclass(frozen=True)
class Run:
model: str
total_steps: int
global_batch: int
dp_workers: int
seed: int = 0
optim: Optim = field(default_factory=Optim)
def __post_init__(self):
if self.global_batch % self.dp_workers:
raise ValueError('global_batch must divide evenly by dp_workers')
if self.optim.warmup_steps >= self.total_steps:
raise ValueError('warmup_steps must be less than total_steps')
if not 0 < self.optim.lr < 1:
raise ValueError('optim.lr is out of range')
def deep_merge(base: dict, over: dict) -> dict:
out = dict(base)
for key, value in over.items():
if isinstance(value, dict) and isinstance(out.get(key), dict):
out[key] = deep_merge(out[key], value)
else:
out[key] = value # scalars and lists replace
return out
def apply_override(cfg: dict, dotted: str, raw: str) -> dict:
*parents, leaf = dotted.split('.')
node = cfg
for part in parents:
if part not in node or not isinstance(node[part], dict):
raise KeyError(f'unknown config path: {dotted}')
node = node[part]
if leaf not in node:
raise KeyError(f'unknown config key: {dotted}')
node[leaf] = yaml.safe_load(raw) # 'true' -> bool, '2048' -> int
return cfg
def strict(cls, data: dict):
kinds = {f.name: f.type for f in fields(cls)}
unknown = set(data) - set(kinds)
if unknown:
raise KeyError(f'unknown keys for {cls.__name__}: {sorted(unknown)}')
out = {}
for key, value in data.items():
kind = kinds[key]
if kind in (int, float, str) and not isinstance(value, kind):
if kind is int and isinstance(value, float) and not value.is_integer():
raise TypeError(f'{cls.__name__}.{key} must be an int, got {value!r}')
try:
value = kind(value) # '3e-4' -> 3e-4, 2048.0 -> 2048
except (TypeError, ValueError):
raise TypeError(f'{cls.__name__}.{key} must be {kind.__name__}, got {value!r}')
out[key] = value
return cls(**out)
def load_yaml(path: str) -> dict:
with open(path) as fh:
return yaml.safe_load(fh) or {}
def build(files: list[str], overrides: list[str], root: str = 'configs') -> Run:
cfg: dict = {'optim': asdict(Optim())}
for path in files:
data = load_yaml(path)
for piece in data.pop('compose', []): # pieces first, in order
cfg = deep_merge(cfg, load_yaml(f'{root}/{piece}.yaml'))
cfg = deep_merge(cfg, data) # then the file itself
for item in overrides:
key, raw = item.split('=', 1)
cfg = apply_override(cfg, key, raw)
optim = strict(Optim, cfg.pop('optim'))
return strict(Run, {**cfg, 'optim': optim})
def fingerprint(run: Run) -> str:
blob = json.dumps(asdict(run), sort_keys=True).encode()
return hashlib.sha256(blob).hexdigest()[:16]Walk the interviewer through four choices. Overrides must target existing keys, which catches typos on the command line too. strict rejects unknown keys at every level. It also coerces and checks types, because dataclasses do not, and because PyYAML parses 3e-4 as a string. That bug is common enough to mention unprompted. The fingerprint hashes a sorted JSON dump of the resolved object, so two runs with identical settings share a fingerprint regardless of how they were composed.
In production, a validation library gives better error messages and nested coercion, and a composition library adds config groups, interpolation, and multi-run sweeps. Name the ones you know and say what you would keep from each.
What separates answers: the loader
WeakMerges and hopes
Merges dictionaries and passes them to the trainer. Typos and wrong types reach the cluster.
GoodTyped and strict
Rejects unknown keys and wrong types, checks cross-field rules, and fails in seconds before launch.
StrongStrict, explained, and testable
Also rejects overrides to unknown paths, documents list-replace semantics, fingerprints the resolved config, and includes unit tests for each rule.
6.2 What reproducible can and cannot mean
Before reading on. A researcher relaunches a run from its record on the same hardware. The loss curve differs slightly. Is the system broken?
Not necessarily. Some GPU kernels are nondeterministic: parallel reductions add floating-point numbers in a varying order, and float addition is not associative. A different GPU count changes the reduction layout. Library upgrades change kernels.
State the promise precisely: the same settings, code, image, data, and seeds, with results that match within normal run-to-run noise. Offer a deterministic mode for debugging, which forces deterministic kernels and fixed reduction orders, and say it costs speed. Record hardware layout and library versions, so a difference can be traced to its cause.
6.3 Relaunch and drift
Relaunching a run from last month uses the recorded image and commit. If the code has since changed in ways the old config does not understand, the recorded image still runs the old code, so the run reproduces. If the researcher wants the old config on new code, the relaunch tool resolves the old config against the new schema and shows the differences: new keys with defaults, removed keys, changed defaults. The researcher confirms with --allow-drift, and the new record notes it.
6.4 Shared defaults that change
Before reading on. The team changes the default warmup from 2,000 to 4,000 steps. What happens to runs launched last week?
Nothing, because their records hold resolved values. What does change is the meaning of old experiment files: launching exp/scaling_7b today gives 4,000 warmup steps. Make that visible. train diff against the last run of the same experiment shows the change, and a default change in code review lists the experiments it affects.
6.5 Sweeps
Hyperparameter sweeps generate many runs from one config plus a grid or random search over override values. Each generated run is a normal run with its own record. Store the sweep ID in each record, so results can be grouped, and make the sweep definition itself a versioned file.
6.6 Interpolation and derived values
Before reading on. A researcher wantswarmup_stepsto be 1% oftotal_steps, and the learning rate schedule to end attotal_steps. How do you express that without copying numbers by hand?
Two options, with different risks.
Interpolation in the config. Allow references such as ${total_steps} and simple expressions. The loader resolves them after merging and before validation, and detects cycles. It is convenient, and it hides logic in configs, where it is hard to test.
Derived fields in the schema. Keep configs as plain values, and compute derived values in code: a property on the schema class, such as warmup_steps defaulting to total_steps // 100 when not set. The rule lives in one tested place, and the resolved config records the computed value.
Prefer derived fields for rules that apply everywhere, and allow simple references for one-off wiring. Either way, the resolved config must hold only concrete numbers, so the run record shows what actually ran.
A useful check at build time: print a short summary of the values that most often go wrong, such as effective batch size, tokens per step, total tokens, and peak learning rate, before launch. Researchers catch many mistakes by reading five numbers.
6.7 Validating against the cluster
Before reading on. A config passes every type check, and the job still fails 20 minutes into startup because the model does not fit in GPU memory. Could validation have caught it?
Often yes, with cheap estimates. Schema validation checks types and simple rules; a second stage checks the config against the chosen hardware:
- Memory. Estimate parameters, gradients, optimizer state, and activations per GPU from model size, precision, parallel layout, and micro-batch size. If the estimate exceeds GPU memory with a margin, fail before launch.
- Layout. Tensor, pipeline, and data parallel degrees must multiply to the GPU count, and the number of layers must divide by the pipeline stages.
- Data. The data snapshot must exist and contain enough tokens for the planned steps.
- Quota. The requested GPUs must fit the team's quota, or the job waits in the queue.
For practice: a 7-billion-parameter model with mixed-precision training and an Adam-style optimizer needs about 16 bytes per parameter for weights, gradients, and optimizer states: 112 GB before activations. On 80 GB GPUs, that must be sharded across at least two GPUs even before activations, so a config that asks for pure data parallelism with no sharding fails the check at once.
These checks run in seconds and save the most expensive failure mode in research: a job that waits hours in a queue and then crashes on startup.
What separates answers: validation depth
WeakTypes only
Validates field types and stops there.
GoodTypes and cross-field rules
Adds rules such as batch divisibility and warmup bounds.
StrongChecks against reality
Also estimates memory per GPU, checks the parallel layout against the GPU count, confirms the data snapshot and quota, and prints the key derived numbers before launch.
- Variants
7.1 Configs in Python
Some teams write configs as Python functions. They are powerful: loops, conditionals, imports. They are also hard to diff, hard to validate statically, and they can run arbitrary code at load time. A common compromise: Python builders that must return plain data, validated by the same schema and recorded in resolved form.
7.2 Secrets and environment
Never put secrets in configs; they end up in run records and logs. Reference them by name, such as wandb_key: ${secret:wandb}, and resolve them at launch into environment variables that are never recorded.
7.3 Configs for serving
The same system can configure inference deployments. The difference is rollout: a serving config change goes through canary and rollback, the way code does.
7.4 At ten times the runs
With thousands of runs a week, the run store becomes the team's research memory. Index run records by experiment, fingerprint, and key settings, so researchers can ask which runs used a given data mix or learning rate. Two runs with the same fingerprint signal duplicated work. Retention matters too: keep every record, but expire checkpoints of abandoned runs.
- The transferable pattern
A configuration system is a compiler with a strict front end and a frozen output: many readable inputs, one validated artifact, and a record that pins everything the artifact depends on. The same pattern runs build systems, infrastructure as code, and feature-flag platforms. Validate early, freeze the output, and make the frozen output the only thing that executes.
Review: the 30-second answer
- Compose from pieces in a fixed order. Later wins; dictionaries merge; lists replace.
- Dotted overrides to existing keys only. Typos fail.
- Typed schemas with cross-field rules. Errors in seconds, before any GPU starts.
- Record the resolved config with commit, image, data, and seeds. The trainer reads only the record.
- Promise equivalence, not bitwise identity. Offer a deterministic mode with its cost.
Quiz
+Why reject unknown keys instead of ignoring them?
Because an ignored typo silently falls back to a default. A run can then burn thousands of GPU-hours on the wrong setting before anyone notices. Rejecting it costs milliseconds at launch.
+Why should the trainer read the run record instead of the config files?
The files can change after launch, through edits or new defaults. The record holds the resolved values the run actually used, so the trainer, a relaunch, and anyone auditing the run all see the same thing.
+What five things does reproducing a run require?
The resolved config, the exact code commit, the container image digest, the data snapshot, and every random seed.
+Why can a relaunch differ slightly even with all five recorded?
Some GPU kernels are nondeterministic: parallel reductions add floating-point numbers in varying order, and float addition is not associative. Different GPU counts and library versions change results further.
+Why do lists replace instead of merge?
Merging lists item by item has no single sensible meaning. Should a shorter list truncate, or should items append? Replacement is predictable, and predictability is the point of a merge order.
+Where should derived values such as warmup steps be computed?
Preferably in the schema, as a tested rule in code, with the computed value written into the resolved config. The run record then shows the actual number that ran.
+Why estimate GPU memory at validation time?
A job that does not fit crashes at startup, often after waiting hours in a queue. A seconds-long estimate from model size, precision, and layout catches it before launch.
Sources and further reading
- Hydra documents config groups, composition, and command-line overrides.
- OmegaConf covers structured configs, merging, and interpolation.
- PyTorch reproducibility notes explain nondeterministic operations and deterministic modes.
- Deployment covers versioning and rollback, the production side of the same discipline.
