On September 26, I celebrated Petrov Day in Cusco, Peru. Our book club, De Profundis, turned one that night, and I asked for five minutes to speak. Most people in the room had never heard of Stanislav Petrov.The 5:26 video is in Spanish with English subtitles.I told the story from memory, in Spanish, with a…
Læs hele historien hos lesswrong.com →On “reversed stupidity”, the success of deep learning, and what a mistaken forecast should change about a model of intelligenceThis began as a Twitter/X thread after I posted a 2007 passage in which Eliezer Yudkowsky was scathing about neural networks, and asked whether, in hindsight, his dismissal was itself a case of “reversed stupidity is…
Læs hele historien hos lesswrong.com →Epistemic status: I don't know much about AI alignment. This is just me thinking out loud.AI use: grammar and minor fixes.Reading about Inkhaven inspired me to try writing something to see how much I enjoy it. I've heard the idea that humans are not aligned in the same sense AI is not aligned. I am…
Læs hele historien hos lesswrong.com →Disclaimer: I am still relatively new to decision theory, and I don't want to come across like this hasn't been discussed extensively. However, I realized that much of my optimism on AI benevolence was based on this line of thinking, so I wanted to formalize it for critique. The literature I could find did not…
Læs hele historien hos lesswrong.com →TL;DR – More of the public than ever are talking about AI consciousness and they won't wait for research to split into their pro- and anti- consciousness camps. The pro-consciousness camp could threaten the control paradigm by viewing monitoring and training as violations, while the anti-consciousness camp could over-invest in control paradigms and miss opportunities…
Læs hele historien hos lesswrong.com →It’s tempting to define safety research as research that enables developers to deploy an AI system more safely without making the deployment much more expensive or much less useful.You can visualize this definition of safety research as pushing out the safety-usefulness Pareto frontier. At any given level of usefulness, there's greater safety available.Awkwardly, this definition counts…
Læs hele historien hos lesswrong.com →This is a beginner-friendly video. Second half might contain new content for people who already know about acausal trade.The beginning introduces decision theory and acausal interactions (CDT, EDT, ECL, MBAT).The middle explains why I think influencing how AIs reason about acausal interactions is time-sensitive.The last bit talks about how to do the influencing.DiscussLæs hele historien…
Læs hele historien hos lesswrong.com →100% human-written. Copyedited by Claude.*Epistemic status: experience report of one ~60h project with Claude Code, plus longer for the write-up.Scope: I supervised Claude's thinking but didn't look at the code. Spec was an exploratory, high-level design refined iteratively.Method: I introduced runtime and design self-checks as Claude failure modes appeared, and ran manual regression tests against…
Læs hele historien hos lesswrong.com →The US has about 150 million workers. How long did it take them to learn to do what they do? I'm interested in this because of the analogous question for AI. Running an AI on a job and training it to do the job are different costs. In a recent post I found that when…
Læs hele historien hos lesswrong.com →TL;DR: We identify a confound in no-cot-bench related to the positioning of the prompt's "key." Correcting for this decreases GPT-6.1 Sol's no-cot reasoning depth by 16%, with the effect likely growing as dependent depth increases. This matters because current benchmark performance reflects both serial reasoning depth and the ability to spread computation over tokens. These…
Læs hele historien hos lesswrong.com →