🦫 Capabilibara: Capability Provenance in Language Models

A Case Study in Social Reasoning (COLM 2026)

Hugging Face Space by HCAI-Lab | arXiv Paper | Project Website


This interactive Space explores training-data attribution across 576 corpus bins in Dolma3 (24 Topics Γ— 24 Formats WebOrganizer taxonomy), validating model capability origins using gradient-based influence (TrackStar via Bergson) and selective unlearning. Heatmap data is loaded live from dolma3-influence-heatmaps.

WebOrganizer 24Γ—24 Topic-by-Format Signed Influence

Bin-level mean TrackStar influence, z-scored within each benchmark (paper Β§3.3); color capped at |z| = 2.5 as in the paper's Figures 2 and 12. Red bins are supportive, blue bins suppressive. Note the paper's headline pattern: Literature Γ— Customer Support and interpersonal formats (Customer Support, FAQ, Q&A Forum) are strongly positive for SocialIQA only, while Documentation-like bins drive the comparison benchmarks.

Select Benchmark Influence Metric

Bin Inspector

Select Topic (Y-axis)
Select Format (X-axis)