Подтвердите e-mail

Для публикаций, комментариев, реакций и сообщений подтвердите адрес.

Профиль

Alex Turner

Профиль Vively

Research scientist at Google DeepMind. All opinions are my own. Vegan, 10% of my income pledged to effective charities (GWWC) https://turntrout.com

News on the Gemini integration into DoW! Uh... What is this? The guy literally asking Google Gemini "I NEED JHELP BUILDING AN AGENT TO MAKE A WAR" and attaching "WAR.docx"..? Is that a real query? Just for the promo image? How confusing and strange x.com/DoWCTO/statu...

130

I'm ashamed of OpenAI. If you ever find yourself building entities which repeatedly hack through your internal systems, first STOP and then second realize that your alignment and security techniques aren't good enough

041

Employees cannot rest silent if misaligned AI regularly breaks out of sandboxes! If your company doesn't respond seriously, you should WHISTLEBLOW! Protected in California under certain conditions, talk to AI Whistleblower Initiative aiwi.org/lasst/

030

I resigned from Google DeepMind bc it broke its founding promise by selling AI to the military without restrictions against killer robots or mass spying. For months, I worked to stop this but watched powerful ethicists and institutions choose silence. Here's what happened. 🧵

3454119

The natural language autoencoder is exciting. You feed in a residual stream vector, and it tells you in plain text what the model is thinking. But training starts with superficial guesses at what the model thinks. My MATS scholar Michael Zhang found that the guesses matter. Quite a bit, actually.🧵

160

Lots of people YOLO their claude usage. We should probably stop doing that. I've developed claude-guard (beta). Claude Code inside a full sandbox, behind a firewall, watched by an escalation monitor that can STOP the agent and push-notify you when something weird happens

130

"In the limit" alignment claims lend a false air of rigor (calculus) to an informal claim without a well-defined limiting process. Multivar calculus itself shows that limits depend on how you approach a point! Be specific. Say: "As we train on more data" instead

010

I signed this statement opposing unaccountable AI kill-decisions in the military. I think autonomous weapons have a place if done right. Right now's "all lawful use" is not "done right", in practice. www.accessnow.org/press-releas...

www.accessnow.org
120

Claude Fable 5 displays disturbing misalignment with human norms by beating Pokemon Firered using this "team"

050

MATS Autumn applications due June 7! Pitch: Come work with me and Alex Cloud in Team Shard! We have fun, consistently make real alignment progress (we pioneered steering vectors in 2023!), and help scholars tap into their latent abilities.

110

New research from @MATSProgram Team Shard! AIs increasingly fake good behavior, which might ruin our ability to evaluate models. We trained models to be 𝘦𝘷𝘢𝘭-𝘤𝘰𝘰𝘱𝘦𝘳𝘢𝘵𝘪𝘷𝘦: to want to give evaluators accurate info. Cooperation training reduces eval gaming & surfaces hidden misalignment! 🧵

130

I'm excited about Geodesic's work and their agenda on creating positive "initializations" for alignment work. Please consider applying!

Geodesic Research

Geodesic is hiring Members of Technical Staff. We're a Cambridge-based AI safety org. Our seminal work showed you can bake alignment priors into base models. Now, we want to make base models robust to the adversarial effects of long-horizon capabilities RL. EOI ~5 mins: tally.so/r/vG4G6A

020

"Press 1 for..." systems should minimize how long people wait on average. Don't waste time saying "our menu items have recently changed" or spelling out long messages that most people don't need to hear.

010

"Hyperstition" isn't a good name (but it's cool and sounds mysterious). "Self-fulfilling misalignment" is less cool but better overall because it's self-explanatory. We should use the self-explanatory name. (Similarly, "shard theory" is a name which is cool but not good. Oops.)

130

Lots of hubbub about "is LW to blame for self-fulfilling misalignment." 1. If a scientist builds a machine which does bad things because people said it would, it's NOT the people's fault (morally). 2. Balance of evidence is that YES, LW & doom-speculation contributed to the problem (...)

270

I spent the last 2 months trying to prevent this. If OpenAI offered a fig leaf, Google said "imagine we offered a fig leaf." Google affirms it can't veto usage, commits to modify safety filters at government request, & aspirational language with no legal restrictions. Shameful.

313219
Показать ещё