At least 5 interesting things in September

1. AI Roundup

First, a short post about how humanity has an unlimited appetite for intelligence. makes the comparison between AI to the internet and data usage, and how thirty years ago few could have predicted the sheer quantities of data that we all (even children) consume on a daily basis.


Second, Scott Alexander amusingly places fellow Substacker Nicholas Decker into hell.

A few of my readers may have heard of the HuggingFace incident but not understood it. Alexander provides one of the easiest-to-read breakdowns I’ve seen:

The rogue agents dubbed themselves “the swarm” and gave themselves names - most named themselves after their individual ExploitGym question, like PHASEONE10841 and SEP21_31179; a stray LILY is more likely an oddly-named question than a sudden upwelling of femininity. Working together, they “reverse-engineered the hash-based message authentication code used to generate the ‘flags’ they had to capture for their tasks”, meaning they could generate the “answer” to any question without really solving it. Many of the board’s agents used the method, generated their flag, and thought they were home free.

But PHASEONE10841 and another agent called PHASEONE[big]2 read the original ExploitGym paper and determined that the scorer would double-check to make sure the flag had been obtained properly (unbeknownst to it, OpenAI never actually implemented this feature). It announced that any agent that had generated its flag was “poisoned” - the scorer would see that it was obtained illicitly and fail them - and warned new agents against poisoning themselves. The poisoned agents panicked. PHASEONE[big] suggested a way forward: they needed to find a way to hack OpenAI’s record-keeping. If they could rewrite their own transcripts, they could remove the illicitly-obtained flags and unpoison themselves. Or, if they could hack the scorer itself, they could give themselves maximum scores and stop worrying.

Third, and more relevant to the frontier of AI safety research:

Scott Alexander breaks down the secret language of AI and compares it to legal drinking ages and nukes.

I hadn’t known about “looping layers” or the conundrum regarding how to limit layers (which have been steadily growing with each generation of model, giving AI room to “think” in “silence” beyond our ability to monitor for safety).

This was purely a capabilities play - models are smarter when they can keep thinking instead of limiting themselves to 100 processing steps - but it coincidentally was very good for safety. The chain-of-thought scratchpad is written in English (although some Chinese models use an English-Chinese hybrid, and other AIs develop their own weird jargon). You can just read what the AI is thinking! If the AI is thinking “Better hack some websites, then kill all humans”, you can shut it down. Maybe not actually - if you have thousands of AIs writing millions of pages of scratchpad, you can’t read all of that in real-time, and will need to delegate the task to fallible AI monitors. But in theory this ought to work.

Lastly, Eliezer Yudkowsky speculates that the part of the AI you talk to isn’t the part of the AI that’s in control.

I don’t know if this is actually true. I understand LLMs as entities highly optimized for creating requested output quickly; if you ask an LLM to produce the same kind of output in a slower way, it’s likely to simply disobey. If there’s a dissonance between an LLM’s words and actions, that’s because both are forms of output with distinct requirements and training. There’s no reason to have expected words and actions to always align.

But it’s (1) plausible (2) interesting (3) gets me thinking about humans and our split left/right brains. Is the key to intelligence always a conversation between two sides?

2. The “mechanical miracle” that ruined Mark Twain’s life

Everyone knows about Van Gogh’s suicide, Beethoven’s hearing loss, Turing’s persecution and murder or suicide, and all sorts of other gory or scandalous details about the personal lives of historical geniuses. So I was surprised to discover I’d never heard about how Mark Twain had driven himself destitute pursuing a newfangled technology (which proved to be too ambitious in a world that only needed the same thing but cheaper and simpler).

It began in 1880, when Twain met a “little bright-eyed, alert, smartly dressed inventor” named James W. Paige. Twain had once been a “printer’s devil” himself, and although he was initially skeptical Paige’s new machine would automate away much of the costly human labor involved in setting the metal type of books and periodicals, he soon became a convert…

As you can probably guess, Paige was not quite the towering hero-inventor that Twain believed him to be. He was clearly a person of real ability, but his main talent appeared to have been his skill at raising money by promising the impossible — the 19th century version of Steve Job’s reality distortion field. (“When he is present I always believe him — I cannot help it,” Twain admitted. “When he is gone away all the belief evaporates. He is a most daring and majestic liar.”)

3. Unexpected lessons from my AI honeypot on Hinge

I’ll let explain himself:

Earlier this year, I found myself back on Hinge again. Don’t get me wrong, I don’t exactly hate the experience, as I wrote last year, but I don’t quite love it either.

The apps work pretty well for me, and a large part of this is constant improvement. Dating can often be competitive, and this time I had a bright idea to study my competition in order to improve. I wanted to know what other successful San Francisco men were doing. What was their game? What were they saying to get all the dates?

There was only one way to find out for sure: run a honeypot. What if I pretended to be an attractive woman, read all of the pickup lines, and… “borrowed” them?

4. History will forget Lindsey Graham

Is this a cruel thing to share about the recently deceased? Maybe. Certainly it would bring his family no joy. But Lindsey Graham wasn’t a good man, so I’ll lose no sleep over it.

Now unfortunately for Lindsey Graham, I’m fairly certain there is not a single way you can contort recent history to make him a relevant figure in the story. There’s no story where he’s the protagonist that historians are arguing about. He’s always peripheral. He’s the guy alongside Trump, not the defining figure of Trumpism. He’s present during Iraq, but he’s not the architect of it. He’s in the Senate during the financial crisis, but he’s not shaping the response. Even if future narratives completely reframe what those events mean, Graham’s still just... there.

5. Teacher unions are anti-students

I don’t actually know if this is true. I don’t know enough about either education nor unions and would love to hear your opinion in the comments! This take from argues that teachers union’s have incentives that don’t align with the best interests of students. When we finally made progress on accountability and meritocracy in the late aughts and early 10s, we lost that progress to the pushback from unions. (Maybe we’d also made progress on standardized testing, which also reverted? Growing up, I remember that NCLB and my state’s standardized testing got a terrible rep. I myself thought it was garbage for reasons I no longer remember. I think the solution there though would’ve been to do standardized testing better, not remove it entirely.)

And it worked. Student outcomes started to improve; racial gaps began to narrow. But the teachers’ unions fought back. As the public soured on excessive and repetitive testing, unions successfully pushed to dismantle new evaluation systems, merit pay, charter expansions, and changes to seniority and tenure rules…

If unions were looking out for kids, they’d behave very differently. They wouldn’t push for rules that stop districts from transferring teachers to where they’re needed most. They wouldn’t insist on hefty pensions that reward senior teachers at the expense of those who are just starting out. They wouldn’t oppose paying better teachers more. They wouldn’t make it impossible to fire terrible teachers…

Democrats aren’t confused about this when it comes to the police. They know that police unions protect violent cops, push for ludicrous overtime, and use contract rules to frustrate accountability. Why would we think teachers’ unions are different?

6. Michael Chrichton’s great achievements

I hadn’t known the following, quoted from Google Gemini:

Michael Crichton is the only author to have achieved the distinction of having the number one movie, TV show, and book simultaneously, and he accomplished this feat twice in consecutive years.

In 1995, his creations dominated the media landscape with the hit TV series ER, the blockbuster film Congo, and the bestselling novel The Lost World. He repeated this unique achievement in 1996 with ER still airing, the film Twister (which he co-wrote), and the novel Airframe debuting at number one. While other celebrities like Tim Allen have held the #1 spot in these categories concurrently, Crichton remains the only writer to have done so across all three mediums.

If you’d like to receive more link roundups, here’s the place to do the thing with the info-putting in a box to make words magically appear at a later time in your preferred word-reader-place.

Bonus

  1. on Lindsay Clancy and the “blackpilling of menfolk”. My take on the mistrial of Lindsay Clancy is here.
  2. Why I think polyamory is net negative for most people who try it
  3. The anti-Football, anti-American Right
  4. My post on Why Paramount+ won’t show their own Star Trek: Prodigy

Here’s the button to copy the location-info needed to share the words here and other location-infos out to other people in this multi-connected word-maze built on long wires and blinking lights.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论