AI & Voice

The Interface Ladder

Guest interfaces have climbed a ladder for thirty years: website, chatbot, voice, the live call. Every rung fixed a flaw and introduced a new one. The top rung, what the guest hears and sees moving in lockstep, is still empty.

Listen
The Interface Ladder

A mother of three, 9:40 on a Sunday night, planning spring break. The hotel’s website is open on her laptop. The hotel’s reservations line is pressed against her ear. The agent on the phone says the Garden King sleeps four and has a partial ocean view. She scrolls the photo gallery trying to find it. She asks if it is the one with the balcony in the third photo. It is not. That is the Terrace King, which is sold out for her dates. They spend four minutes discovering they have been talking about two different rooms.

Look closely at what she is doing. She is running two interfaces at once, one audio and one visual, and performing the synchronization between them herself, in her head, at 9:40 on a Sunday, for free. The hotel’s most motivated customer of the evening is working as unpaid middleware between the hotel’s own voice channel and the hotel’s own screen.

A few weeks ago we argued that the next year of voice AI will be won by the harness, not the model. Buried in that piece was a single clause: one Operator sub-agent exists “to manage what the guest sees on screen if there is a screen.” That clause deserves its own essay. The screen is not a detail. It is the next rung of a ladder the industry has been climbing for thirty years without naming it.

We are going to name it: the Interface Ladder.

Every rung of the ladder fixed one flaw and introduced another

Line up every interface a hotel has ever offered a guest and you get a sequence, ordered not by chronology but by how much of a real conversation each one can carry.

Rung one is the static website. It fixed the busy signal and the brochure. Always on, visually rich, and photography sells rooms in a way no sentence ever has. The flaw it introduced: the guest does all of the work. Search, scroll, compare, infer, and when the site cannot answer the one question that matters, whether the Garden King’s sofa bed actually fits a nine-year-old, she picks up the phone anyway.

Rung two is the chatbot. It fixed navigation. Ask instead of hunt. The flaw it introduced is the one the entire industry politely ignored for a decade: chat is a terrible container for a travel decision. Walls of text describing things that should be shown. Typing on a phone at night. No way to hold three options side by side.

Rung three is the voice agent. It fixed the keyboard. Speaking is faster than typing, listening is easier than reading, and the conversation finally moves at a human pace. The flaw it introduced is blindness. You cannot read a photo gallery aloud. Asking a guest to compare three room categories by voice alone is a working-memory test she did not sign up for, and the transcripts we graded for the harness piece are full of the moment she fails it: “wait, which one had the balcony again?”

Rung four is the live phone call with a human. Still the interface guests reach for when the stakes are real, because a good reservations agent brings judgment, empathy, and improvisation no rung below can match. But the blindness is still there: agent and guest are still describing pictures to each other. And the rung does not scale. Hold times, hours of operation, a labor line.

Notice the irony. Rung four is the oldest interface on the ladder. Thirty years of software climbed away from the phone call, then climbed back toward it, because the call carried something every intermediate rung dropped: conversational bandwidth. The ladder is not a march of progress. It is a search, and it is not finished.

Because the top rung is still empty. The interface where what the guest hears and what the guest sees move together, in lockstep, in the same conversation. Say “show me the ocean view rooms” and the rooms appear as the words land. Say “the second one” and the second one opens. The cart builds itself on screen while the voice confirms it.

We have all seen this interface. Cinema has been rendering it for decades: the bridge of a starship, the lab in an Iron Man film, a person talking to a computer while something visual answers in the same beat. Nobody in a movie talks to a progress spinner or reads a wall of text back to the machine. The audio-visual interface is so obviously the destination that film treated it as a default fifty years ago. Obvious in hindsight, brutally hard to build. That combination is exactly what leaves a rung empty this long.

Brian Chesky has already conceded that the top rung exists

If you want a second opinion on the ladder, take it from the most design-literate CEO in travel. On Airbnb’s Q1 2026 earnings call on May 7, 2026, Brian Chesky laid out what he described as four problems with chatbots for travel: they produce too much text, they offer no direct manipulation, they are bad at comparison, and travel bookings are multiplayer while chatbots are stubbornly single-player. His formulation of the fix is an interface that stays conversational but is far more visually rich than a chatbot. And he said the part executives usually skip: it is a hard problem, and nobody, Airbnb included, has cleanly solved it.

Four weeks later, on June 4, 2026, Bloomberg and TechCrunch reported that Chesky is standing up a new AI lab, separate from Airbnb, focused on user interaction and interfaces beyond the chatbot.

Read the sequence plainly. The CEO who rebuilt his product around photography looked at the chatbot era and said: wrong container. Then he pointed at the empty rung, conversational plus visual, admitted nobody has cleanly reached it, and is building a lab to go after it. The claim that the top of the ladder is empty is not ours alone. It is Airbnb’s, on the record.

Everyone is climbing toward the same empty rung from a different wall

Map the field honestly and you find three camps, each holding one piece of the answer and missing another.

The generative UI camp is adding rich visuals to text. Google shipped Dynamic View and generative interfaces with Gemini 3, where the model composes a bespoke visual layout for each answer instead of a paragraph. Thesys C1 sells an API that turns LLM output into live interfaces. Vercel’s AI SDK generative UI tooling and OpenAI’s Apps SDK let developers render real components inside a conversation. The thinking layer is further ahead than the shipping layer: Geoffrey Litt’s Malleable Software in the Age of LLMs, Maggie Appleton’s essays on language-model interfaces, and Linus Lee’s work on generative interfaces beyond chat all converge on the same conclusion: chat was the demo, not the destination. But nearly all of it is anchored to typed conversation. Visuals bolted onto text.

The OTA camp shipped the best visual-chat agent so far, and it is silent. On June 3, 2026, Priceline relaunched Penny as a fully agentic assistant, running a coordinated system of more than ten specialized agents, with Anthropic’s Claude driving the core reasoning, and a live interactive map that updates as the conversation moves. It is, as far as we can tell, the closest thing to lockstep any large travel player has shipped at scale. And it is text. There is no voice on the rung Penny stands on.

The hotel voice camp answers the phone with nothing to look at. Vendors like Canary have put AI voice on hotel phone lines, and the category is real and growing. But the guest’s screen plays no part in the call. It is rung three, executed well.

Camp                      Has                        Missing
------------------------  -------------------------  --------------------
Generative UI platforms   Rich visuals on text       Voice
Priceline Penny           Agents + live map on chat  Voice
Hotel voice vendors       Voice on the phone line    Any screen at all
Top rung                  Voice + visual, lockstep   (empty)

Three camps, one empty rung, and every camp is exactly one layer short of it.

The word doing all the work is lockstep

Why has nobody assembled the two halves? Because voice and screens are physically different media, and the difference bites at the integration seam.

Voice is serial, ephemeral, real-time: one word at a time, gone as it lands, interruptible mid-syllable. A screen is parallel, persistent, stateful: many things at once, held until told otherwise. Lockstep means both channels express the same state at the same moment, every moment: the room the voice describes is the room the screen highlights, a barge-in halts the render as well as the speech, and the cart on screen is not a picture of the booking state, it is the booking state.

Teams that try to bolt these together usually build a second system that listens to the transcript and guesses what to draw. That architecture drifts, because the guesser and the speaker share no state, and drift in this interface is fatal. A chatbot that renders the wrong card is a bug. A voice saying “the Garden King” while the screen highlights the Terrace King recreates, pixel for pixel, the Sunday-night failure at the top of this essay, except now the hotel built it on purpose.

This is why we wrote the harness piece first. A voice agent harness is the runtime layer that owns the call’s state: which tools are mounted, which observations are live, which agent is operating. Once a system genuinely owns call state, the screen stops being a rendering problem and becomes what it should have been all along: a second actuator driven by the same state machine as the dialogue. Voice and visuals cannot drift when they are two outputs of one brain. Without a harness, lockstep is a heroic synchronization project. With one, it is a mounting decision.

We have shipped one early, imperfect version of the top rung

FlowStay’s current product puts a guest in a voice conversation, on the phone or on the property’s website, while a synchronized surface shows what the conversation is about: the rooms with their photos, the options under discussion, the cart assembling itself as the guest and the agent agree. The screen is driven by the same harness state that drives the voice, which is the only reason the two stay honest with each other.

We want to be precise about what this is and is not. It is early. The layouts are narrower than we want, the choreography between utterance and render still has rough edges, and there are conversational branches where the screen goes conservative and shows less rather than risk showing wrong. It is one implementation, not the definition of the category.

And it is a window, not a moat. Airbnb has a dedicated lab and the deepest design bench in travel. Priceline is one voice channel away from putting Penny on this rung. Any generative UI platform could add an audio loop. The honest description of our position is a head start measured in production calls and harness maturity, not a wall anyone will bounce off. Windows close. That is what makes them windows. Our plan is the same one we stated in the harness piece: publish the numbers, build in public, and be standing on the rung, with paying hotels, when the crowd arrives.

The guest was never supposed to be the middleware

Back to the mother of three. It is still Sunday night, but run the scene on the top rung. One surface. She says the dates and the party size, and the available rooms appear as the agent speaks. She says “the one with the balcony,” and it opens, and the agent says the sofa bed fits a nine-year-old because the agent and the screen are reading the same state. The cart builds while they talk. Eleven minutes, booked, done, and at no point did she hold two disconnected interfaces together with her own attention.

She closes the laptop. The phone was the laptop. The laptop was the phone.

One guest. One conversation. One surface.

Not chat. Not voice alone. Both, in lockstep.

Sources

  1. Airbnb (ABNB) Q1 2026 Earnings Call Transcript (May 7, 2026) The Motley Fool
  2. Airbnb CEO Brian Chesky plans a new AI lab focused on user interaction (June 4, 2026) TechCrunch / Bloomberg
  3. Priceline relaunches Penny as a fully agentic travel assistant (June 3, 2026) Priceline Press Room
  4. Canary AI Voice Canary Technologies
  5. Gemini 3: Dynamic View and generative interfaces in the Gemini app Google
  6. Thesys C1: Generative UI API Thesys
  7. OpenAI Apps SDK OpenAI Developers
  8. Vercel AI SDK: Generative User Interfaces Vercel
  9. Malleable Software in the Age of LLMs Geoffrey Litt
  10. Maggie Appleton: Essays on interfaces and language models Maggie Appleton
  11. Linus Lee: Generative interfaces beyond chat Linus Lee
← Back to all posts Book a demo →