Skip to content
← Back to writings
AI Safety

Blue Dot Technical AI Safety – Course Reflection

August 2026

This course follows my completion of Blue Dot Impact AGI Strategy – Course Reflection.

🤖AI Transparency Statement: This article was fully written by me, then went through basic grammar checks and a very light editing feedback loop on clarity with Claude Opus 5.0.

Course Overview

The Technical AI Safety course aims to give a broad overview of the technical tools and methods we have to achieve AI safety. It starts off with an overview of the challenges, then moves through the pipeline from pre-training, training, mechanistic interpretability (understanding how models operate and why they make the decisions they do), evaluations, through to control and monitoring. It closes with envisioning how you could use one of those methods to defend against one of the key harm categories, and then envisioning your next steps in the space. Full details, including the entire curriculum, are available here.

My Experience

I entered the course with a tentative focus on automating evals after sharing my first writing piece on them, and left with that focus unchanged, but with an expansion on some ideas that I had initially shared around democratizing and standardizing the research, data, standards and more. I deepened my knowledge in several areas beyond evals, especially mechanistic interpretability. I left the course feeling validated in my direction, and with a stronger focus. I’m still envisioning an entry for a project and potentially the first role being in automated evaluations, but I’m starting to know more about where my skills and experience can benefit the space the most. I’m thinking about how I can zoom out, and help support the bigger picture through efforts to standardize, democratize and build user experiences based on my years of experience doing those things.

My New Working Hypothesis

In absorbing all of this, I started to find a new hypothesis that I want to keep testing:

The AI safety space is very fragmented and at an early stage of maturity while evolving rapidly. Different labs run different evals and benchmarks in different tools, present their findings differently, as well as their policies. Making a decision or assessment on whether a model is safe for release or exhibiting dangerous capabilities continues to fall back to human judgement. It feels as though we’re reaching an inflection point now where we do need to start adding some connective tissue between the frontier labs, independent orgs and researchers, and the government. I see this need on at least two levels, although they do relate:

  1. The research writeups, data, code, methods, prompts, results, etc need to be unified into a single experience for anyone to review, search, and to perform meta-analysis on to determine the state of the space and known gaps whether that’s research, standardized practices, current model capabilities, latest benchmarks and evals, etc.
  2. We need a centralized body to bring everyone together, align on standards and best practices, develop them and work on having them broadly adopted and implemented correctly. This isn’t a regulatory or licensing body. They can’t levy punishments themselves. It’s an opt-in that everyone contributes to make the space safer faster through combining their strengths and generating standardized work that can be highly trusted.

I’ll continue to validate these ideas through conversations as I join projects, events like EAGx, or other web events like evals paper reading clubs. There’s also a lot to be gleaned from reading between the lines of forum posts (AI Alignment Forum, Less Wrong, EA Forums, etc), shared research and other announcements.

What I Learned

  • Anthropic and others publish Responsible Scaling Policies (RSPs), model cards and related docs. But, our observation was that the evals don’t seem aligned and consistent. Everyone is doing their own thing.
    • Further, once I got into the details, it seems like the thresholds to stop an unsafe deployment are very squishy and can be overridden. This reflects the pressure the frontier labs are under to release as fast as possible in the race with each other.
  • It’s not just the RSPs or Frontier Labs either. When it comes to evals, benchmarks, etc everyone is working in their silos, posting and sharing in a myriad of places. This fed my above hypothesis.
  • Early in the class, I took a position on mechanistic interpretability (mech interp) being the most important tool. My reasoning was that if we can understand how AI works, why it’s making decisions, and when certain things happen, it would alleviate a lot of the downstream pressures of trying to detect misbehavior through evals, control and monitoring. I did learn a lot about mech interp with one of the most interesting things being linear probes to objectively observe and log when certain behaviors are happening, including in evals. By the end of class my position evolved back to mech interp being one important tool among many, partly due to the state of mixed results and impact we’re getting from it right now.
  • Sometimes the best option, unfortunately, is going to really strong defenses, and some of those might be AI aided. This also reaffirmed that technical AI safety as prevention is not the only solve:
    • There will be untrusted models out there, no matter what.
    • When it comes to risks like gradual disempowerment, the biggest levers come down to informing the general public and forcing policy around it.
  • It’s mind boggling how much of this ecosystem already relies on AI running on training filters, evals, monitoring, etc. I get it, when it comes to scalable oversight, we just can’t get enough people in. There’s also a matter of efficiency of dollars spent, especially on simpler tasks.

Other Notes

  • The “future” was happening during this course. This course took place after the Mythos escape and controversial release of Fable earlier this spring and summer. Then during the course the OpenAI HuggingFace incident broke, followed by Anthropic and Meta confirming similar (Reuters news article as I didn’t see an official Meta post on this). This absolutely cements the need for these techniques and beyond to ensure we have AI that is as aligned as we can make it, followed by strong controls.
  • I still think we need more people in evals, and that they don’t need to all be highly trained and top-notch researchers. I have a piece backlogged for this.
  • Some of the biggest value for my role and next steps came from the final day content, especially this Less Wrong piece on iterators, collectors and amplifiers. It helped define the org-size (medium, growing), their challenges, and defined what role an amplifier (my archetype) would play in that. This was an important signal on my fit in the space, helps me narrow my focus and how to position myself outwardly.
  • There’s a trove of valuable information in the optional resources too. I paid more attention to it than I did in AGI Strategy. I do have a list of optional resources and exercises that I want to follow up on.

Feedback on the Course

I would overall very strongly recommend this to anyone interested in the technical AI safety space. The content is well curated and relevant. It’s great to work through it with a cohort. I like the 1-1 discussions paired with group discussions.

  • I think we all had the feedback that reading took more time than expected and estimates should be adjusted in the course info and curriculum.
  • One of my criticisms of the AGI Strategy course was that some content was a bit dated, especially in the advances of agents doing autonomous work. The content here felt more fresh and up to date.
  • I would say most of the content is geared towards future AI safety researchers and engineers. While I didn’t struggle to keep up with it from a technical standpoint (a positive sign for me), I would say it would be nice to see more of a throughline through the course of how generalists are contributing or playing a role also.

What’s Next for Me (Fall 2026)

Last Update: August 28th, 2026 (may be quite out of date by the time you read this)

Update: I went to EAGxBerkeley 2026 and evolved my plan and assessment further here.

Original Text:

Overall the plan is to keep learning, growing, connecting with the community and contributing to AI safety through generating new outputs.

There’s an abbreviated version of most of the content in this post in my updated 1-pager that supersedes the one I shared in Blue Dot Impact AGI Strategy – Course Reflection

  • Projects: Applied to Fall 2026 Projects – SPAR and the next BlueDot Technical AI Safety Project Sprint – This should give me hands-on experience with credentialed projects and mentor, while also producing more public outputs, and ultimately making my first contributions to the space.
  • Courses: Self-Pacing ARENA ch0 and ch3 to grow Python, PyTorch and evals knowledge – Will help answer common questions on applications regarding programming and evals experience
  • Community: Joining EAGx, BlueDot Evals Reading Club, and local meetups – Helps me learn about how the space is evolving and its needs, and connects me to potential next opportunities.