𦫠Capabilibara: Capability Provenance in Language Models
A Case Study in Social Reasoning (COLM 2026)
Hugging Face Space by HCAI-Lab | arXiv Paper | Project Website
This interactive Space explores training-data attribution across 576 corpus bins in Dolma3 (24 Topics Γ 24 Formats taxonomy), validating model capability origins using gradient-based influence (TrackStar) and selective unlearning.
WebOrganizer 24Γ24 Topic-by-Format Taxonomy Matrix
Select Benchmark Influence Metric
Bin Inspector
Select Topic (Y-axis)
Select Format (X-axis)
Headline Study Scale & Key Results
| Metric | Value | Detail |
|---|---|---|
| Corpus Bins | 576 |
WebOrganizer 24Γ24 topic-format matrix |
| Working Set | 5.68M |
Stratified unique Dolma3 documents |
| Base Model | OLMo-3-7B |
AllenAI open base model |
| Unlearning Shift | +1.60 pp |
SocialIQA damage on unlearning flagged bins ($p \approx 10^{-5}$) |
| Attribution Compute | ~37K |
H200-equivalent GPU hours |
Key Findings
- Social reasoning vs. STEM provenance diverge: Social reasoning capabilities depend strongly on informal discussion, personal essays, and Q&A formats, whereas STEM capabilities concentrate in technical manuals and academic papers.
- Unlearning validation: Targeted unlearning on top-attributed bins significantly degrades target capabilities while leaving un-targeted capabilities intact.
Cite This Work
@inproceedings{matlin2026capabilityprovenance,
title = {Capability Provenance in Language Models: A Case Study in Social Reasoning},
author = {Glenn Matlin and Chandreyi Chakraborty and Saehee Eom and Mika Okamoto and
Rayan Castilla and Louis Jaburi and Alvin Deng and Taywon Min and
Lucia Quirke and Stella Biderman and Mark Riedl},
booktitle = {Proceedings of the Conference on Language Modeling (COLM 2026)},
year = {2026},
eprint = {2606.19625},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2606.19625}
}