On my first full day back in San Francisco, I had dinner in a busy hotel lobby and spent the evening exploring dot on my Mac.

I laughed. I got up for a short walk. At one point I said, “That's effing crazy.”

The interesting part was how easily I could keep adding thoughts while something was already happening.

OpenAI announced dots on September 29. The official guide describes an agent that can keep working between conversations, use connected apps, and bring results back for review.

This is a field report from one evening. I was testing the interaction, following my curiosity, and occasionally trying to break the flow.

Sending before the thought is finished

I started sending tiny messages.

“also”

“and reply”

Sometimes a correction. Sometimes another question before the previous work had finished.

These were submitted messages. I wasn't testing whether dot could see unsent keystrokes. I wanted to know whether I could express a stream of consciousness without first packaging it into one complete prompt.

It incorporated the additions and corrections into its responses. I could keep shaping the request while work was underway.

While dot researched public reactions, I could scroll back, read earlier replies, and add questions. Reading, asking, and waiting overlapped. I didn't have to organize the whole experience around the next answer arriving.

I kept coming back to the word “blurring.” The boundary between giving input and watching work happen felt less rigid.

The observable behavior is straightforward: I sent more information, and later responses reflected it. That doesn't tell me whether some hidden process was continuously thinking. It tells me the interaction accommodated an unfinished thought.

Following the conversation around

I also tested replies under individual messages. I replied in a branch, returned to the main conversation, then returned to the branch. The reply targets stayed intact.

That sounds small. It matters when several questions are alive at once. A correction needs to belong to the right question, especially if other work is still running.

Voice exposed a different set of boundaries. I tried interrupting and asked dot to speak faster, including at twice the speed. I didn't measure whether it reached an exact multiplier.

When I switched apps on my iPad, a call dropped. That was one observation, with no diagnosis. I also asked for transcripts because they hadn't appeared automatically. Dot pasted the text it had received; the final portion was cut off when the call ended.

Those details matter to me. If voice is part of an ongoing workspace, I want to know which words made it through and which didn't.

I liked the interface, too: the fade and the muted blue gradient. They contributed to the feeling of the evening. I wouldn't treat that feeling as evidence of capability, but it belongs in a report about using the product.

Recovering context, with gaps

Another test was looking backward.

Dot recovered a screenshot from a 2023 plugin hackathon. It also produced a partial index of conversations from July through October. Some titles and links were missing.

Recovering something specific was useful. The gaps were useful evidence, too. I couldn't treat that index as a complete record of my history, or assume that a relevant detail would always be available later.

We also looked at two projects, Shared API and MandateGuard. The inspection was read-only. Shared API had immutable request construction; MandateGuard had a policy scaffold. No tests were run.

That gave us something concrete to discuss without turning a code read into a claim that either project worked end to end.

In Artifact a Node, I wrote about an idea surviving pressure testing and becoming available for future retrieval. This evening produced a candidate: useful agent interaction should let me revise my intent while preserving a clear account of what the system is doing.

Separating the layers

There are at least three things to evaluate here: the model, the agent runtime, and the interface.

The model produces judgments and responses. The runtime coordinates work, state, tools, and permissions. The interface determines how I can inspect and influence that work.

My strongest observation was about the combined interaction. A single evening can't isolate how much came from each layer. It also can't establish a leap in Astra's intelligence or long-term reliability.

Asynchronous agent work already exists. What caught my attention was experiencing it through a conversation I could keep contributing to.

Three tests I want to run next

The next artifact could be a small, steerable coding-task agent. We discussed it; we didn't build it.

I want to test three things:

  1. A correction that arrives late. Start a task, change a requirement while it is running, and inspect the result. Does the agent identify which work became stale? Does the final artifact reflect the correction, or merely acknowledge it in chat?

  2. Stop and resume. Interrupt a multi-step task, then resume it. Check the recorded state, completed steps, and remaining work. Can I tell what happened while I was away, and does resuming avoid duplicate actions?

  3. Approval of an exact diff. Ask for a code change, inspect the diff, and approve only that version. If the proposed change grows or changes afterward, require another review. A smooth conversation still needs a precise boundary around execution.

My workflow remains chat, specification, Git, Codex, tests and review, then commit and push. This article is one small instance of that process.

I left the evening excited about a specific possibility. I could think in fragments, redirect work, and return to earlier questions without rebuilding the entire request each time.

Now I want to see whether that feeling survives a task where a missed correction actually changes the result.