2 Sources
[1]
Guide Labs debuts a new kind of interpretable LLM | TechCrunch
The challenge of wrangling a deep learning model is often understanding why it does what it does: Whether it's xAI's repeated struggle sessions to fine-tune Grok's odd politics, ChatGPT's struggles with sycophancy, or run-of-the-mill hallucinations, plumbing through a neural network with billions
[2]
New Steerling-8B model can trace every single word back to its training source
Guide Labs, a San Francisco-based startup, announced the open sourcing of Steerling-8B, an 8-billion-parameter large language model. Co-founded by CEO Julius Adebayo and Chief Science Officer Aya Abdelsalam Ismail, the company introduced the model on Monday. The architecture enables full
Share
Copy Link
San Francisco startup Guide Labs has open-sourced Steerling-8B, an 8-billion parameter language model built with a novel architecture that makes every generated token traceable to its training data origins. Founded by Julius Adebayo and Aya Abdelsalam Ismail, the company aims to solve AI's black box problem by engineering interpretability directly into the model rather than analyzing it post-hoc.
Guide Labs, a San Francisco startup founded by CEO Julius Adebayo and Chief Science Officer Aya Abdelsalam Ismail, has open-sourced Steerling-8B, an interpretable LLM that fundamentally changes how developers understand what their AI systems are doing
1
. The 8-billion parameter model was released on Monday with a novel architecture designed to address one of AI's most persistent challenges: understanding why deep learning models make the decisions they do. From xAI's struggles to fine-tune Grok's political leanings to ChatGPT's issues with hallucinations, the black box problem has plagued AI developers for years1
.The breakthrough lies in Steerling-8B's ability to trace every token produced by the model back to its origins in the training data
1
. This capability ranges from simple tasks like verifying reference materials for cited facts to complex analyses of how the model encodes abstract concepts like humor or gender1
2
.
Source: TechCrunch
Adebayo explained the complexity: "If I have a trillion ways to encode gender, and I encode it in 1 billion of the 1 trillion things that I have, you have to make sure you find all those 1 billion things that I've encoded, and then you have to be able to reliably turn that on, turn them off"
1
. While this is technically possible with current models, the process remains fragile and unreliable.The architecture fundamentally alters the standard transformer structure by inserting a concept layer that categorizes data into traceable buckets during training
2
. "The kind of interpretability people do is...neuroscience on a model, and we flip that," Adebayo told TechCrunch. "What we do is actually engineer the model from the ground up so that you don't need to do neuroscience"1
. This approach requires more upfront data annotation, but Guide Labs used other AI systems to assist in the labeling process1
2
.Adebayo began this work during his doctoral studies at MIT, co-authoring a widely cited 2018 paper that demonstrated the unreliability of existing methods for understanding deep learning models
1
2
. One concern with engineering interpretability is whether it might eliminate emergent behaviors that make language models valuable. Adebayo says Steerling-8B still exhibits these capabilities, with his team tracking "discovered concepts" that the model generates autonomously without explicit training, such as quantum computing1
2
.Related Stories
Guide Labs positions the technology as essential for high-stakes applications requiring strict control over LLM outputs
2
. In consumer-facing AI systems, the architecture enables developers to block copyright infringement and better manage outputs around sensitive subjects like violence or drug abuse1
2
. For regulated industries, particularly finance, the model can evaluate loan applicants based on financial records while explicitly excluding protected attributes like race, addressing regulatory requirements1
2
. In scientific research, including protein folding, the model provides insight into why specific predictions succeed, addressing a critical gap in computational biology1
2
.Steerling-8B achieves approximately 90% of the capability of existing frontier models while using less training data, thanks to its novel architecture
1
2
. Adebayo argues that "training interpretable models is no longer a sort of science; it's now an engineering problem," suggesting the approach can scale to match frontier models with significantly higher parameter counts1
2
. The company emerged from Y Combinator and raised $9 million in seed funding led by Initialized Capital in November 20241
2
. Next steps include building a larger model and offering API access and agentic capabilities to users1
2
. "As we're going after these models that are going to be super intelligent, you don't want something to be making decisions on your behalf that's sort of mysterious to you," Adebayo said1
.Summarized by
Navi
13 Jan 2026•Science and Research
19 Feb 2026•Science and Research

31 Jan 2025•Technology

1
Science and Research

2
Policy and Regulation

3
Technology