Skip to content

Discover a circuit.
Then act on it.

CircuitKIT is a unified framework for mechanistic interpretability.
13 discovery algorithms · 6-pillar faithfulness evaluation · real interventions

Get started View on GitHub

LSAL v1.2 Python 3.10+ 13 algorithms 6 pillars

Given a model and a task, CircuitKIT discovers the circuit driving that behaviour, evaluates how faithful it is, and lets you act on it — prune, quantize, edit, steer, or fine-tune — then export a reloadable HuggingFace checkpoint.

  • 13 discovery algorithms


    EAP family, ACDC, IBCircuit and CD-T, each with an explicit stability tier. Six are stable (eap, eap-ig, eap-gp, acdc, ibcircuit, cdt) and have been tested across the GPT-2, Llama, Gemma, and Qwen families. The other seven are research: validated on GPT-2/IOI but not yet exercised at scale or across architectures. Default: eap-ig.

    Algorithm overview

  • 6-pillar faithfulness


    Causal patching, ablation, stability, robustness, baselines, generalization. Six complementary checks, not one score.

    Evaluation framework

  • Real interventions


    Structural pruning, circuit-aware quantization, ROME/MEMIT editing, activation steering, circuit-restricted LoRA. Writes a reloadable HuggingFace checkpoint.

    Applications guide

  • Modern models


    Works on instruction-tuned Llama-3, Gemma, Qwen — not just GPT-2. GQA/RoPE, chat templates, grouped-query attention handled natively.

    Supported models

See it in 60 seconds

The model field is the only thing that changes across models — each tab below uses a different one (gpt2 on CPU, Qwen2.5 and Llama-3.2 on GPU/MPS) to show the same workflow is model-agnostic.

from circuitkit.api import discover_circuit, evaluate_circuit

circuit = discover_circuit({
    "model": {"name": "gpt2", "precision": "float32"},   # CPU-friendly default
    "discovery": {"algorithm": "eap-ig", "task": "ioi",
                  "data_params": {"num_examples": 32}},
    "pruning": {"target_sparsity": 0.3, "scope": "both"},
    "output_path": "./circuit.pt",
})
results = evaluate_circuit({
    "model": {"name": "gpt2"},
    "discovery": {"algorithm": "eap-ig", "task": "ioi"},
    "pruning": {"target_sparsity": 0.3, "scope": "both"},
    "output_path": "./circuit.pt",
})
from circuitkit import Pipeline

# any HF model — Qwen 2.5 / 3, Llama 3, Gemma 2 / 3, Pythia, GPT-2 ...
pipe = Pipeline("Qwen/Qwen2.5-0.5B-Instruct", task="ioi")
pipe.discover(algorithm="eap-ig", n_examples=128, sparsity=0.3)
pipe.evaluate(pillars=["patching", "ablation", "baselines"])
pipe.prune(sparsity=0.3)
pipe.export("./output/checkpoint")
pipe.summary()
# Llama & Gemma are gated — accept the license on HF, or swap for an open
# model like Qwen/Qwen2.5-1.5B-Instruct. (This block prunes, so use a
# registered arch — Pythia is discovery/eval only and cannot prune/export.)
circuitkit discover --model meta-llama/Llama-3.2-1B-Instruct --algorithm eap-ig \
    --task ioi --sparsity 0.3 --level node --output ./circuit.pt
circuitkit evaluate --model meta-llama/Llama-3.2-1B-Instruct --artifact ./circuit.pt
circuitkit prune --model meta-llama/Llama-3.2-1B-Instruct --artifact ./circuit.pt \
    --sparsity 0.3 --output ./pruned
circuitkit benchmark --models gpt2 Qwen/Qwen2.5-1.5B-Instruct --tasks ioi \
    --algorithms eap-ig --interventions prune

How it works

flowchart LR
    A["Model + Task<br/>(GPT-2, Llama, Gemma, Qwen...)"]
    A --> B{"discover_circuit({...})"}

    subgraph backends["13 algorithms / 4 backends"]
        direction TB
        B1["EAP family  Stable<br/>eap, eap-ig, eap-gp"]
        B2["ACDC / IBCircuit / CD-T  Stable"]
        B3["RelP, PEAP, ...  Research"]
    end
    B --> B1 & B2 & B3

    B1 & B2 & B3 --> C[("Circuit artifact<br/>circuit.pt + scores")]

    C --> D{"6-pillar faithfulness"}
    D --> D1["1 Causal patching · 2 Ablation<br/>3 Stability · 4 Robustness<br/>5 Baselines · 6 Generalization"]

    D1 --> E{"Act on it"}
    E --> E1["Prune"]
    E --> E2["Quantize"]
    E --> E3["Edit"]
    E --> E4["Steer"]
    E --> E5["Fine-tune"]

    E1 & E2 & E3 & E4 & E5 --> F["HF checkpoint →<br/>reload & benchmark"]

Tutorials & case studies

Start here Open Hardware
01 Quickstart Pipeline — discover → evaluate → prune → export Colab no GPU
02 Algorithm Comparison — 6 algorithms head-to-head Colab GPU helps
03 Evaluation Deep Dive — all 6 faithfulness pillars Colab GPU helps
23 Jailbreak Refusal Localization — find vs. act, across 3 instruction-tuned models Colab GPU (T4+)

All 8 tutorial notebooks, the 13 numbered scripts, and 11 domain case studies (compliance auditing, banking safety, gender-bias mitigation, permanent unlearning, edge deployment) are catalogued in the Examples overview.

Where next

Getting started Install, quickstart, core concepts
Guides Per-topic usage guides with code
Trust & Audit Stability tiers, audit status
Reference API reference, CLI reference
Examples Scripts, notebooks, and case studies

Cite

@software{circuitkit2026,
  title  = {CircuitKIT: Circuit Discovery, Evaluation, and Application Toolkit
            for Mechanistic Interpretability},
  author = {Seth, Pratinav and Gosalia, Hem and Kasliwal, Aditya
            and Sankarapu, Vinay Kumar},
  year   = {2026},
  url    = {https://github.com/Lexsi-Labs/circuitkit}
}