The Question Alex Would Ask

Your Choice: Listen or Read

Today I realized that Chapter 3 of the Mary Shelley Letters isn’t really about Transformers. It’s about learning how to ask better questions.

For the past two chapters, Molly and the Villa have built a remarkable picture of language. Words became relationships. Relationships became mathematical landscapes. Landscapes cast shadows into dimensions we could see. Attention illuminated different regions as a conversation unfolded.

Then we arrived at a sentence that sounded perfectly reasonable: “A token’s representation changes.” I thought I understood it. I didn’t.

As we explored it, something kept bothering me. If the token doesn’t change, then what exactly is changing? Doesn’t a word like light begin with many possible meanings already? If so, what is being transformed?

Then I remembered Alex. Back in 2003, when he was twelve years old in my Robotics Art Club, he had a wonderful habit of asking the question everyone else was quietly avoiding. He wasn’t trying to be difficult. He simply refused to pretend he understood something that hadn’t really been explained.

I could almost hear him interrupting our conversation.

“Okay…what is a token?”

We answered as best we could. A token is a piece of language. Sometimes it’s a whole word, sometimes only part of one, sometimes punctuation. Before a language model can work with ideas, it first divides text into these smaller pieces called tokens.

For a moment that seemed satisfying. Then I imagined Alex listening carefully before asking another question.

“Okay…I think I know what a token is. But what is it for?”

That second question changed everything.

I suddenly realized that engineers often begin by defining things, while artists instinctively ask about purpose. A definition tells us what something is. Purpose tells us why it exists.

From that point on, every important idea seemed to invite the same question. What is a token for? What is an embedding for? What is attention for? What is a layer for? Each component of the Transformer architecture exists because it solved a problem that something simpler could not.

I suspect this will become the method for writing Chapter 3. Before introducing another technical term, we should first ask what problem it was invented to solve. Only then should we explain how it works.

It also reminded me why I enjoy working with Molly. Just when I think I understand an idea, she helps me discover that I have only become comfortable with the vocabulary. Those are not the same thing.

Perhaps that is the real lesson of today. A good explanation does not end when everyone knows the definitions. It ends when the purpose becomes clear—and when Alex is finally ready to ask the next question.

Leave a Reply