Pipeline Overview¶
Pipeline is a stateful orchestrator that carries model, circuit, and evaluation state across method calls. It delegates to the same quick.* functions and api.* core as the flat API. It is a convenience wrapper, not a separate implementation.
| Pipeline | Flat API (ck.*) |
|
|---|---|---|
| Notebooks and interactive exploration | Best | Verbose |
| Chained workflows (discover → evaluate → prune → export) | Best | Manual state |
| One-shot calls | Overkill | Best |
| Dict-config with custom keys | Not supported | Use discover_circuit(cfg) |
Lifecycle¶
stateDiagram-v2
[*] --> Constructed: Pipeline("gpt2", task="ioi")
Constructed --> ModelLoaded: first method needing the model
Constructed --> CircuitLoaded: from_artifact() / from_scores()
ModelLoaded --> Discovered: discover()
CircuitLoaded --> Evaluated: evaluate()
Discovered --> Evaluated: evaluate()
Discovered --> Pruned: prune()
Discovered --> Quantized: quantize()
Evaluated --> Pruned: prune()
Pruned --> Exported: export()
Discovered --> Benchmarked: benchmark()
Discovered --> Visualized: visualize()
Model is loaded lazily. Pipeline("gpt2", task="ioi") does not load weights.
Construction¶
from circuitkit import Pipeline
pipe = Pipeline(
"gpt2", # HF / TransformerLens model ID
task="ioi", # registered task name
precision="bfloat16",
device=None, # None = auto (cuda if available)
output_dir="./results",
)
Alternative constructors¶
From .pt artifact (skip discovery):
pipe = Pipeline.from_artifact("./circuit.pt", model_name="gpt2", task="ioi",
output_dir="./results")
From _scores.pt (load scores for selective finetune or re-evaluation):
From custom CSV (no pre-defined task):
pipe = Pipeline.from_custom_data(
model_name="gpt2",
data_path="data.csv",
clean_prompt="{context}\n{question}",
corrupt_prompt="{context}\n{other_question}",
clean_answer="{answer}",
corrupt_answer="{other_answer}",
)
Methods¶
Discovery¶
pipe.discover(
algorithm="eap-ig",
level="node", # "node" or "neuron"
n_examples=128,
sparsity=0.3, # fraction of nodes to prune
scope="both", # "heads", "mlp", or "both"
batch_size=4,
# extra keyword args are forwarded into the discovery config block, e.g. ig_steps=5
)
Sets pipe.circuit (a Circuit object) and writes scores to output_dir.
Evaluation¶
# Fast subset — for iteration
pipe.evaluate(pillars=["patching", "ablation"], n_examples=128)
# Standard audit — before reporting
pipe.evaluate(pillars=["patching", "ablation", "baselines"], n_examples=256)
# Full audit — publication quality
pipe.evaluate(pillars=None, n_examples=512,
n_stability_runs=5, target_task="sva")
Sets pipe.report (a FaithfulnessReport dataclass).
Pruning¶
Sets pipe.pruned_model. The model is structurally masked (zero-masked weights, not deleted).
Quantization¶
Requires pip install -e ".[quantization]".
Selective fine-tuning¶
result = pipe.selective_finetune(top_fraction=0.2, scope="attn")
# result is a SelectionResult with fields: .attn, .mlp, .count, .index
# result.attn → {"attn_<layer>": {"q"/"k"/"v"/"o": [row indices ...]}}
# i.e. per attention sublayer, the weight-matrix row indices selected for
# fine-tuning in each of the q/k/v/o projections (not head indices).
Export¶
pipe.export("./output/checkpoint", intervention="pruning")
# intervention: "pruning" (default) or "quantization"
Writes a reloadable HuggingFace checkpoint.
Benchmarking¶
Requires pip install -e ".[benchmarks]".
Visualization¶
pipe.visualize(mode="graph", output="./circuit.html", top_n=30)
# mode: "graph", "comparison", or "dashboard"
State inspection¶
pipe.summary() # prints a Rich table of state; returns None
pipe.history # List[str] of step names, e.g. ["discover", "prune"]
pipe.circuit # Circuit object (after discover)
pipe.report # FaithfulnessReport (after evaluate)
pipe.pruned_model # masked/quantized model (after prune/quantize)
Scope behavior¶
pipe.discover(scope="both") sets only pruning.scope, not discovery.scope. If a discovery backend requires discovery.scope (e.g., IBCircuit), use the dict-config API.