📰 Key Takeaways

Baseten teamed up with Hugging Face and Goodfire AI on Wednesday to roll out a new safety infrastructure standard for open-weight models, launching its research division Base Labs at the same time. The move comes as concerns over open-weight model safety are heating up—a technique called “abliteration” can strip a model’s built-in safeguards, making it dangerous. The scale of the problem isn’t small: over 6,000 abliterated models are already listed on Hugging Face, the platform that hosts open-source models.

Base Labs will focus on developing and publishing methods for training and monitoring open models, positioning this as a “standard” for open models—one where safety is built into training and deployment from the start, not bolted on afterward. Baseten said on X: “We believe openness is an advantage for AI safety. Openness gives us more transparency into model behavior, and most importantly, it lets us turn safety research into controls that are actually enforceable and auditable—something closed models can’t offer.”

The three companies haven’t disclosed specific technical details of the partnership yet, but Goodfire, in a reply post, framed the goal as: “Safety needs to be built into open models, and it’s on the people serving them to make that happen.” Goodfire, which specializes in reverse-engineering AI’s “black box” to explain how models make decisions, looks like the natural fit for handling the “built-in” part of that equation.

Baseten closed a $1.5 billion Series D this June at a $13 billion valuation. Goodfire AI is similarly well-funded, having raised a $150 million Series B led by B Capital earlier this year to push forward its model interpretability platform. Baseten also put out an open invitation to the broader developer ecosystem to contribute to the framework.


💬 JudyAI Lab Take

Baseten teaming up with Hugging Face and Goodfire AI to launch Base Labs is aimed at closing a real gap in open-weight model safety—over 6,000 abliterated models, stripped of their safeguards, are already circulating on Hugging Face.

We see this as a sign the industry is shifting. The risks of open weights used to get waved off as “just the cost of openness.” Now a few well-capitalized players (Baseten at a $13B valuation, Goodfire fresh off a $150M Series B) are moving to bake safety directly into training and deployment, instead of patching it on after the fact. Goodfire’s job is decoding how models actually make decisions; Baseten’s job is the deployment infrastructure. Together, they’re turning “safety” into an enforceable technical spec rather than just a talking point. It also says something about where the open-model ecosystem needs to go—community self-policing alone isn’t cutting it, and someone has to set the standard first.

Next time you’re picking an open-source model, it’s worth checking whether it’s been abliterated—make that a basic part of your safety checklist.


📅 Source Info


🔗 Further Reading