TL;DRi) ML4Good is an intensive eight-day technical AI safety bootcamp covering topics such as AI agents, alignment, model reasoning, evaluations, interpretability, and AI governance.ii) It provides useful information about fellowships, career opportunities, starting an AI safety startup, and the broader AI safety ecosystem.iii) I would encourage people interested in AI safety to apply, including those…
Læs hele historien hos lesswrong.com →Cross-posted from Substack.Below, I’ve included 5 recommendations for hiring managers in AI Safety and 5 tips for candidates as they navigate their career search. I’ve also highlighted orgs I think are doing hiring particularly well (from a candidate’s perspective) and strategies to avoid.Note for Readers: I do not have a technical background, and have never…
Læs hele historien hos lesswrong.com →SummarySynthetic document finetuning (SDF) — the state of the art for implanting false beliefs into LLMs — produces training documents with features that distinguish them from pretraining text.These features (i.e. synthetic markers) are partially responsible for making SDF false facts distinguishable from pretraining-acquired beliefs by linear probes on middle-layer activations.Reducing synthetic markers enables SDF to…
Læs hele historien hos lesswrong.com →TL;DRIn Q2 2026, we ran the first MATS x Coefficient Giving (CG) Pitch Week: a four-week program designed to take fellowship researchers from an early idea to a funding decision. The program was a joint effort between MATS, CG, Constellation, and Catalyze Impact.38 expressions of interest; all were invited to a two-week Exploration Sprint. 17…
Læs hele historien hos lesswrong.com →tl;drIf you are working in AI safety, while presenting information to interested non-AI safety or transitioning personnel, avoid portraying high values of p(doom) without sharing counterarguments. In addition, please give them time and point them towards emotional support so the unprecedented risk can be processed at their own pace. Encourage those transitioning and working in…
Læs hele historien hos lesswrong.com →Epistemic status: personal reflections on my life and worldviews.AI use: review, grammar and readability changes.I discovered Inkhaven recently, and it got me interested in writing. Partly I want to produce posts for the application, and partly – why wait until it starts when I can just start writing now? This made me think about why…
Læs hele historien hos lesswrong.com →AI safety has gone mainstream. More people than ever want to work on AI safety research, and there are real policy wins in the US, the UK and Europe. But political will and market incentives are now the bottleneck, and policy officers can't own correct implementation alone.Having built a Responsible AI program at an Accenture…
Læs hele historien hos lesswrong.com →From a distant vantage point, the Tamed Animal Farm looked like any other. It had the farmhouse, the barns, the windmill, the hayfields, the garden, the barbed-wire fence, the woodpiles, and the pickup truck. Yet zoom in on the farm and you'd find something odd, the animals weren't cows and pigs, they were coyotes, raccoons,…
Læs hele historien hos lesswrong.com →IntroductionWe continue where we left off. We will now analyze the special case when the true distribution is regular for the statistical model. This covers the classical Bayesian statistics. We analyze the regular posterior distribution, and observe the universal structure that is the asymptotic behavior of this distribution. The theorem is akin to central limit…
Læs hele historien hos lesswrong.com →I have seen plenty of discourse and writing on the subject of intent vs impact when actions lead to harm, particularly in existing interpersonal relationships. The thrust tends to be that it's unreasonable to focus the discussion or conclusions on the intent of the actor rather than the impact the actions had on the other…
Læs hele historien hos lesswrong.com →