Подтвердите e-mail

Для публикаций, комментариев, реакций и сообщений подтвердите адрес.

Профиль

Juan Diego Rodriguez

Профиль Vively

CS PhD student at UT Austin in #NLP Interested in language, reasoning, semantics and cognitive science. One day we'll have more efficient, interpretable and robust models! Other interests: math, philosophy, cinema https://www.juandiego-rodriguez.com/

It turns out the "13-step-ladder plan" meant "evaluate the 13 checkpoints" lol

040

I find Fable loves using jargon and fortune-cookie talk...

100

And other people’s papers too!

000

I really don’t get it. There are so many good reasons to stay away from it! It is no longer Twitter — it is a far-right propaganda machine. Now I’m forced to at least maintain some presence, because so many other researchers are active there.

000

I wonder how many people put the things the reviewers might object to in the footnotes (or appendices)

110

Sometimes I wish I could add footnotes to my footnotes

220

Right, there are people like Euler who could do math even when blind (the rest of us have to use pencil and paper to get anywhere.). But I think most humans need tools for most things.

110

But humans have always used tools. Not sure what human-level means without tools

100

I wish there were a better term for LLM+harness than "AI agent"

110

I think the honest summary is: they did better than saying nothing, and worse than having the controls that would have made the blog post unnecessary. Whether that's "enough" is partly a policy question about what standards we want to hold labs to, and we don't really have those standards yet."

000

Hoping to start this weekend!

010

Claude is finding good upper bounds while GPT 5.6 breeds slime molds to make longer snakes

020

I got Opus to push it down to 232. It is probably lower than that... Ran out of credits for the week anyway.

020

Wow this is cool! Thanks for posting

110

and use them to achieve permanent military superiority or perpetrate incredibly deep repression of their own people...”

010

The document still rings true if you swap some entities around: “My primary concern is the risk that authoritarian governments—not solely the United States Government (USG), although the USG is clearly the most capable threat—build AI models that are more powerful than those built by China,

141
Показать ещё