At a glance
Peter Yang thinks the interesting thing about Grok Bot is not the model, it is the machine: a persistent computer in the cloud, with its own browser and its own operating system, that stays logged into your apps while your laptop is closed. He builds five bots live on that machine and shows every prompt: an Advisor that invents and spawns the other bots, a YouTube Researcher that mails him a morning brief of outlier videos and comment themes, an X Scout that mines his own timeline and bookmarks for the ten posts he actually wanted to see, a Marie Kondo that audits Gmail, Google Drive and his recurring charges and then unsubscribes, trashes and cancels on his approval, and a Personal Concierge that reads his Japan trip document and watches flight prices. The single most valuable result in the video is a boring one: the concierge found that a Tokyo round trip is about $2,700 cheaper than the open jaw itinerary he had planned, which a plain price alert would never have surfaced because it would never have questioned the route. A bonus round has the agent install Doom, Red Alert and Commander Keen on its own machine, which half works and is more revealing than if it had worked cleanly. He closes with a verdict he does not soften: the biggest obstacle is trust, not capability, and at $200 a month against ChatGPT at $20, Grok Bot is not his daily driver yet.
What a "dedicated cloud computer" actually means (0:41)
The video opens with the claim in one line: Grok Bot is the future of personal agents. Then he backs up and explains why he went looking for it in the first place.
A few weeks before filming, he tweeted that ChatGPT was one feature away from being the AI product he actually wants. That missing feature is not a better model, a longer context window, or another reasoning mode. It is a dedicated personal computer in the cloud for running AI agents. Shortly after the tweet, Lee, from the team building Grok Bot, sent him a DM asking him to test it.
So what is a dedicated cloud computer, and why does it matter enough to build a product around?
He answers with the meme. You have seen people walking around with a laptop half open, lid propped, because closing it would kill the agent mid task. That posture is the whole problem stated as body language: the agent is a process on your machine, so your machine has to stay awake, plugged in, and carried around for as long as the work takes.
"You don't have to do that anymore if you're using Grok Bot." (1:21)
Grok Bot lives on a computer sitting somewhere in xAI's server fleet rather than in your house. It is a real machine, not a sandboxed function call. It has its own browser and its own operating system, and, crucially, you can log into your favorite apps on it and stay logged in. That last property is the one everything else in the video depends on. Every bot he builds is really a bot plus a session cookie that nobody logs out of.
The three way comparison: local agent, cloud browser, cloud computer
He puts three options side by side, and this is the part worth understanding before any of the bots make sense.
Hermes on a Mac Mini at home. This is what he has been running. It sits on 24/7, so the agent genuinely has a persistent computer with persistent logins. The cost is that you have to go buy the machine and set the whole thing up yourself. You can technically run the same setup on a virtual private server instead, but that is more setup work, not less, and now you are also a sysadmin.
ChatGPT work mode. It uses plugins and a cloud browser, so the compute is somebody else's problem. Two things spoil it for him. First, the browser as it stands cannot stay signed in to your favorite apps, which quietly disqualifies most of the automations in this video: a bot that has to ask you to log in again on every run is not an automation, it is a chore with extra steps. Second, the UX is scattered across chat, work mode and Codex. He screen shares it and narrates the mess honestly: a pile of chat threads here, Codex over there, another tab up top for work and chat, and it is confusing even to someone who does this professionally.
Grok Bot. You get a persistent cloud computer out of the box, with no machine to buy and no VPS to configure. Of the three, he calls it the easiest to set up and get going.
The interface, and why it is not a chat log
The second difference is softer and he is upfront that it is a taste argument: the interface feels more focused and more delightful. Each bot is a named thing in a list with its own personality and style, and when a bot starts working it plays a small animation. He demos it by typing a throwaway task, "look up the weather," to show the subtle motion when a bot spins up.
He credits the design team behind the product for that, naming Jenny Wen, whom he has interviewed on this channel before.
"Overall, I think this UX feels more like talking to a coworker than getting lost in a hundred chat threads." (3:10)
That sentence is the real product thesis of the video, and it is worth separating from the marketing. A chat thread is a transcript you have to re-read to know what state you are in. A bot is a persistent named agent with standing instructions, a schedule, and a body of work. The five builds that follow are all the same move: turn a thing you would have re-typed into a chat window every week into a bot that already knows and already ran.
The question he holds open for the rest of the video: can Grok Bot actually replace ChatGPT and Codex, which are his current daily drivers?
Bot one, the Advisor: the bot that builds your other bots (3:30)
His advice is unambiguous. The first bot you should create after installing Grok Bot is an Advisor, and he keeps his pinned to the very top of his bot list.
The job of the Advisor is not to do work. It is to know enough about your work and your life to tell you which bots are worth building, and then to build them. So you start by telling it who you are. Here is what he told his, close to verbatim:
I'm a creator and founder of Behind the Craft. I have two daughters and
I live in the Bay Area. I publish newsletter posts and YouTube videos
focused on practical AI tutorials and podcast interviews. I also have a
membership portal.
Based on all this, suggest five bots that can save me time or money.
For each, explain what it would do in a few sentences.
Two things about that prompt are doing the work. It is specific about the business (a creator business with a newsletter, a channel, podcast interviews and a paid membership) and it is specific about the life (two daughters, the Bay Area), because the bots he wants are not all work bots. And it constrains the output shape: five, with a few sentences each, so the answer is scannable rather than an essay.
The Advisor came back with five ideas:
- Daily AI news scout
- YouTube comment miner
- Membership concierge
- Social media poster
- Podcast prepper
Then comes the move that most people miss, and it is the reason this is bot number one rather than a nice to have:
"What many people don't realize is that you can actually get one bot, in this case our Advisor bot, to help us create the rest of our bots." (4:14)
He does not open a new window and start from scratch. He asks the Advisor to merge two of its own ideas into one bot, in one line: build a single bot that both researches other channels and mines comments, and call it YouTube Researcher. The Advisor creates it and hands it its opening instruction.
Bot two, the YouTube Researcher: outliers and comment mining (4:30)
The Advisor kicked this one off with its own instruction: create a morning brief that sends YouTube intel every morning. Reasonable, and not good enough, so he takes over and gives it specifics.
First instruction: monitor named channels in his niche. He does not ask it to go find interesting YouTube content in general, he points it at a specific competitive set. You watch it start monitoring channels in real time.
The first brief is bad, and he shows it anyway. This is the most useful thirty seconds in the video for anyone who has tried this and given up. The initial report is verbose and hard to follow, a wall of stuff. It did do one thing right: it pulled the comments from his own videos and found themes in them. But as a document it is unusable.
So he writes the format he wants, rather than complaining that the output is bad:
Give me the report in this specific format:
1. My top three content ideas
2. The top five outliers from the other channels I follow
3. The top performing videos overall
4. The top common themes from my own comments
Limit your research to the last 14 days.
The 14 day window is not arbitrary and he explains it: YouTube in AI is intensely topical, so anything older than two weeks is noise for the purpose of deciding what to film next. The definition of outlier is also his, and it is the right one: a video that beat its own channel's typical performance over that window. Not the biggest video, the most over performing one, which is the only version of the metric that transfers to a channel of a different size.
Then the line that generalizes past this bot:
"By the way, this is how you should be working with AI to refine its output, instead of trying to one shot something." (5:52)
The second report is much more concise. Top three content ideas: one is a video he had already planned to make, one is a video on Grok Bot, which is the video you are watching. Then the outliers from other channels, with topics he says are genuinely worth following up on. Then top performers overall, then the common themes from his comments.
Then he schedules it. A daily job sends him the report every morning, and he shows that morning's delivery: top angles, a watch list, and comment themes. What he has bought himself is a research assistant that proactively finds topics to make videos about and feeds back what his own audience is asking for, without him opening anything.
And here he states the pattern he will use for the rest of the video:
- Kick the bot off with an initial prompt.
- Iterate back and forth until the output is actually good.
- Schedule a routine or job (daily, weekly or monthly) so it does the work proactively.
Bot three, the X Scout: mining your own timeline (7:00)
The premise for this one is a data advantage. Grok Bot is part of xAI, so it should have first party access to X data rather than scraping it through a browser like everyone else.
The personal premise is more relatable, and he is candid about it:
"Even though I have an unhealthy addiction to X, I still miss plenty of great content, or bookmark tweets that I never look at again." (7:12)
That is the actual failure mode of the platform for a heavy user. It is not that there is nothing good, it is that the good things arrive unsorted at 3am and the bookmark folder is where interesting posts go to die.
His prompt, close to verbatim:
Find the most viral X posts from the last seven days in my niche and
create a weekly report with the top 10 posts from people I follow,
engaged with, or bookmarked, grouped into a few categories.
Include the full post copy, a link, and your analysis.
End with three content ideas for me to post about.
Note the three sources it draws from: accounts he follows, accounts he has engaged with, and his own bookmarks. That last one is what turns a discovery feed into a personal retrieval system. And note the ending: it does not stop at reading, it converts what it read into three things he could post, which is the job he actually has.
The bot found his X account and automatically created a routine to run this every morning, without being asked.
The first output grouped into three themes:
- Playbooks people save. He calls out one from a creator he follows, Greg, titled "23 ways I use AI agents to grow my startup," plus a post from a Chinese developer on going from loops to graph engineering, and more.
- Grok Bot, Codex and ChatGPT work mode compared and actually used. Posts from Riley, and from Ben on the team behind Grok Bot listing the most loved internal use cases.
- Creative workflows, which he says is a category he genuinely wants to test.
Then the three things to post about next.
And then, because it costs one sentence to ask, he adds a category for fun: the top five funniest posts from his timeline. The number one funniest, as usual, is OpenAI and Anthropic sniping at each other and joking around. The rest are in the same AI niche. This is a small thing that matters more than it looks: the bot is already reading his whole timeline, so the marginal cost of also extracting the funny is zero, and it makes the report something he wants to open.
Getting it out of the app and into his inbox
The most practical moment in this section is when he stops typing and uses voice:
"Can you send this report to my email, and also make sure you include the full post copy in your report. Send it right now." (9:15)
The reason is a real one and it applies to every bot in this video. He does not want to open Grok Bot each morning to see the output. He wants it in the inbox he already opens.
The email arrives with the themes and the full post copy included, with the three things to post about at the bottom. The mechanism is unglamorous: it can send email because he has connected a set of plugins, including the Gmail plugin.
So the loop for X Scout ends up as: initial prompt, have it pull the information, iterate until the report reads right, then either schedule the job inside Grok Bot or push the output into email. He prefers email:
"I like to wake up with the top tweets to consume directly in my inbox instead of having to open a separate app." (10:00)
Bot four, Marie Kondo: the one that actually deletes things (10:30)
This is the bot with the highest stakes and the best payoff, and it is the one he most recommends building. It is named after Marie Kondo, and its job is your digital clutter.
He identifies three places clutter piles up:
- Your email, in the form of newsletters you never open.
- Your Google Drive, in the form of large and abandoned files.
- Your paid subscriptions, in the form of recurring charges you forgot you authorized.
The third is the one with money attached, and it is also the one that is genuinely hard to audit by hand, because the evidence is scattered across years of receipt emails.
He connects both the Gmail and Drive plugins, then gives it this:
Audit my Gmail, Google Drive, and recurring email receipts, and create
a cleanup plan.
- Find newsletters I never open.
- Find large or abandoned Drive files.
- Identify paid subscriptions from my email receipts.
Group everything into categories. Use a numbered list.
Do NOT move, delete, unsubscribe, or cancel anything without my approval.
He stops the video to underline the last line, and he is right to:
"Crucially, I told it to not move, delete, unsubscribe or cancel anything without my approval. The last line is really important for a bot that cleans up your files. You always want to review what it plans to delete before letting it remove anything." (11:12)
He also notes an optional fourth source: you can hook up the Mercury MCP server to pull the charge data directly, Mercury being the bank he uses for his business. That is a strictly better signal than parsing receipt emails, because it is the ledger rather than the paperwork about the ledger.
What actually happened, including the parts that did not work
Grok Bot first confirmed it was connected to Google Drive and Gmail. Then it connected to Mercury over MCP, and here is a detail that tells you what kind of product this is: he had to sign in to Mercury on the remote cloud computer to make that work. Not on his laptop. On the machine in the datacenter. Hold that thought, because it comes back in the closing section.
Then the first report, and it is bad in a specific and familiar way:
"This is a massive list of stuff to clean up. Honestly, this list is pretty overwhelming. It's incredibly long." (11:50)
He gives feedback: list it properly, categorize it properly. Still incredibly long. So he does the thing that works:
Show me a maximum of 10 items in each list, and get rid of all the
random labels.
Now it is digestible: a block of emails, a block of Google Drive files, a block of paid subscriptions. One more round of feedback narrows it to the actionable set only:
Only show me emails to unsubscribe from, Google Drive files to delete,
and paid subscriptions I want to cancel.
That third pass is the important one. The difference between a list of everything in your account and a list of decisions you have to make is the difference between a report and a tool.
The numbered list trick
With the final list in front of him as a numbered list, approving work costs him a sentence:
Email: unsubscribe to 3, 4, 5, 6, 7 and 8.
Google Drive: delete the extra tax return, delete some of these large files.
Paid subscriptions: cancel 11 and 15.
He pulls the general lesson out explicitly, and it is the single most portable tip in the video:
"It's always good to ask Grok Bot or AI to give you its response in a numbered list like this, to make it super easy for you to just tell it to do things by referring to the number in the list instead of having to type everything over again. I use this pattern all the time." (12:40)
The bot takes action
Now it runs, and he calls this the part where the magic happens. It unsubscribes from a batch of senders. It trashes three Google Drive files. Then it starts cancelling subscriptions: Lovable and Equip Foods, a protein company.
Both cancellations required him in the loop, and the friction is worth recording precisely. To cancel Lovable he first had to sign in to the cloud computer with his Lovable credentials, which he did manually, and then they discovered Lovable was already scheduled for cancellation anyway. Then it asked him to sign in to Equip Foods, he did, and the bot cancelled the protein subscription.
He pauses to be fair to the vendor he just cancelled:
"By the way, Equip Foods is a great protein company. I'm only canceling because I have too many protein powders at home already." (13:30)
The measured result:
"It did all this in around five minutes, when it would have taken me probably 30 minutes to an hour to do all this manually." (14:00)
That is the honest number in the video: roughly a 6x to 12x saving on a chore, on a task he was never going to get around to doing by hand.
Making it talk like Marie Kondo
Then the joke that is also a lesson about personality being a feature, not decoration. He decides the bot sounds too much like a robot and not enough like its namesake:
Can you talk like Marie Kondo from now on? Give me an example.
The bot rewrites its own report voice, and the result is the funniest moment in the video, a cleanup log delivered as a small ceremony:
"The files have completed their work. We thank them and place them in the trash, where they may rest. Equip Foods no longer sparks joy. We release the prime protein subscription with gratitude." (14:35)
His closing advice on this bot: set Marie Kondo to run weekly or monthly to clean up your digital files, and never drop the numbered list requirement, so you review before anything is removed.
"You don't want it to accidentally delete some important file." (14:50)
Bot five, the Personal Concierge: the $2,700 sentence (15:00)
This is the bot he wants for all vacation and travel planning, and it produces the single best result in the video.
The setup: he has a vacation document listing his December trip to Japan, a full itinerary he built with AI. He previews it on screen. The flights are not booked yet, so what he wants is price monitoring, but smarter than an alert.
The prompt:
Read my vacation document and monitor the exact flight legs for my
family trip. Get the dates and the best options for each leg.
Check regularly and let me know when the price improves.
It found the flight legs in his document and started checking prices on Google Flights.
He asks the obvious objection out loud before you can: what is the advantage of doing this in Grok Bot instead of just setting a Google Flights price alert? His answer is the point of the whole bot:
"The value here is that Grok Bot can understand my whole trip based on my document and decide what's a better option for my family." (15:55)
An alert watches a route you already chose. This thing read the trip and questioned the route.
The finding
His planned itinerary was an open jaw: SFO to Tokyo, then Tokyo to Fukuoka, which is where the family is going, then Fukuoka back to SFO. It is the shape the document specified and the shape that looks obviously correct when you are planning a trip that ends in a different city from where it starts.
Grok Bot found that a Tokyo round trip is about $2,700 cheaper than that open jaw booking.
"A simple Google price alert would not have found this." (16:35)
That is correct, and it is worth being precise about why. A price alert is a function of a query, and the query is the itinerary. If the itinerary itself is the expensive decision, no amount of monitoring will tell you, because you never asked about the alternative. The agent had the trip document, so the search space it was allowed to consider was "get this family to these places on these dates", not "watch these three legs."
He then asks it to check every morning at 9:00 a.m. to see whether the price improves, and shows a follow up run: the price is still roughly the same, and the Tokyo round trip is still much cheaper.
Where he wants to take it
He is explicit that price watching is the beginning, not the product. Eventually he can ask the bot to go ahead and book the flight, or to check him into the flight when the time comes. And more generally:
"It's always a good idea to have a travel thread or travel bot to both help you plan travel ahead of time, and also, when you're at a location, to help you book amusement parks and figure out what to do every single day." (16:57)
That is the shape of a concierge: one persistent agent that holds the whole trip, before and during, instead of a fresh chat every time a question comes up.
Bonus, the Gamer: what happens when the agent owns a real machine (17:30)
Because Grok Bot comes with a dedicated cloud computer, he does the thing you would do: he asks it to install and let him play retro games. Specifically Red Alert, Doom and Commander Keen.
It found the files and installed them. Then, "open Doom for us to play."
Doom comes up. He opens it inside the virtual cloud computer, starts a new game, picks an episode, picks Hurt Me Plenty, and this is where it falls apart:
"Here's kind of where Grok Bot falls apart a little bit. Because it's on a virtual cloud computer, the mouse isn't quite configured right to actually play Doom. For some reason it's looking at the floor all the time, and I can't seem to adjust the mouse to look up." (18:05)
A first person shooter with a broken vertical axis is not a game, so he moves on to Commander Keen, which he introduces with genuine affection as an awesome platformer he played in his youth. It loads. New game, one player, normal difficulty. He immediately forgets how to jump, discovers that Control is jump, and reports the honest result: it plays better than Doom, but there is enough lag in the keyboard and mouse to make the character difficult to control.
The conclusion is measured and it is not a dunk:
"Grok Bot is not replacing your gaming PC or GeForce NOW yet. But the fact that the agent can install and launch these games on its own computer gives you a sense of how open ended this could become." (19:10)
That is the right read. Nobody needs an agent to play Commander Keen. What the experiment demonstrates is that the machine is a real machine with a real filesystem and a real package situation, and that the agent has enough control over it to download, install and launch arbitrary software without a human touching a terminal. Every serious bot in this video is a consequence of that same capability.
He allows himself one speculation, and flags it as a dream rather than a roadmap: maybe xAI could eventually use its datacenters and GPUs (he jokes, in space) to deliver AAA games through the virtual cloud computer.
A look inside the machine
Then a small moment that is more informative than the games. He opens the file manager on the cloud computer and browses what is there. He notes it looks a bit like Windows 3.1 and admits he is not sure what it actually is. Inside: the Japan flights work from the concierge bot, other files from earlier bots, the games they just installed, plus Chrome and a terminal.
That is the whole thesis made concrete. The outputs of your bots are not messages in a chat log, they are files on a computer, sitting next to a browser and a shell, in a place that is still there tomorrow.
| Bot | Connected to | What it automates | Cadence | What it actually returned |
|---|---|---|---|---|
| Advisor | Nothing. It only needs to know you. | Deciding which bots to build, then creating them | On demand | Five bot ideas, then spawned the YouTube Researcher on request |
| YouTube Researcher | YouTube channels in his niche, his own video comments | Competitive research and audience feedback | Daily brief every morning | Top three content ideas, five channel outliers over 14 days, top performers, comment themes |
| X Scout | X (follows, engagements, bookmarks), Gmail plugin | Reading his own timeline and rescuing dead bookmarks | Daily, delivered to his inbox | Top 10 posts in three themes, full post copy, three ideas to post, plus the five funniest |
| Marie Kondo | Gmail, Google Drive, Mercury over MCP | Unsubscribing, deleting files, cancelling paid subscriptions | Suggested weekly or monthly | Unsubscribed a batch of senders, trashed 3 files, cancelled 2 subscriptions in about 5 minutes |
| Personal Concierge | His vacation document, Google Flights | Monitoring flight prices against the whole trip, not one route | Every morning at 9:00 a.m. | A Tokyo round trip about $2,700 cheaper than the planned open jaw |
| Gamer (bonus) | The cloud computer itself | Installing and launching retro games | On demand | Installed all three. Doom unplayable (mouse look), Commander Keen playable with input lag |
The honest take: trust, privacy and price (20:30)
He saves the hard part for last, and he does not hedge it.
Trust is the bottleneck, not capability
The biggest hurdle for Grok Bot adoption, in his view, is trust. He illustrates it with the smallest possible example, which is why it lands: a Google sign in screen.
"When I see a Google sign in screen like this on my laptop, I don't really think twice before signing in. But because this appeared on the virtual cloud computer, I hesitated a bit, because how do I know that nobody else is seeing this screen on the virtual cloud computer?" (20:30)
That hesitation is the entire product category's problem in one sentence. Every capability in this video, the Gmail audit, the Drive cleanup, the Mercury connection, the Lovable and Equip Foods cancellations, required typing real credentials into a machine he does not own and cannot inspect.
He goes to the Grok Bot website and scrolls to the privacy note at the bottom, which answers the question with the standard set of assurances: it uses the same single sign on and privacy mode you already trust, the cloud computer is encrypted in transit and at rest, and there is no AI training on top of it.
His response to that is the most honest thing in the video, because he does not pretend the assurance settles it:
"I'm willing to give Grok Bot and the virtual cloud computer access to all this stuff because I'm an early AI adopter. But I can see normal people struggling to understand what this cloud computer thing even is, and hesitating to sign into their favorite apps on this device." (21:05)
And then the prediction:
"I think the AI agent platform that figures out trust will be the first to get mass adoption." (21:20)
Worth sitting with. The claim is not that the encryption is insufficient. It is that a correct security posture that a normal person cannot form a mental model of does not produce trust, and trust, not capability, is what gates adoption. Nobody needs to understand TLS to sign into Gmail on their own laptop, because they understand the laptop. Nobody yet understands the cloud computer.
The verdict
On the direction of travel he is unequivocal:
"Grok Bot is a clear sign of the future. We're moving away from manually using our keyboard and mouse to do work on our laptops, to using our voice to orchestrate a bunch of agents that live in a dedicated cloud computer. And Grok Bot is the first product to actually enable this." (21:30)
On whether he is switching, he is equally unequivocal, and the reason is price:
"It's not quite my daily driver yet, because I think ChatGPT still offers more for $20 a month, while Grok Bot requires paying $200 a month to use on a regular basis." (21:45)
That is a 10x price gap for a product he has just spent twenty minutes praising, and he states it without softening. What he does credit it with: a much cleaner UI than ChatGPT right now, and being very capable.
| His own axes | Grok Bot | ChatGPT |
|---|---|---|
| Price for regular use | $200 / month | $20 / month |
| Breadth of what you get | Focused on the cloud computer and bots | "Still offers more" for the money |
| Interface | "Much cleaner UI right now", named bots with personality | Split across chat, work mode and Codex, "kind of a mess" |
| Persistent computer | Yes, its own browser and OS, out of the box | Cloud browser, but no machine of yours |
| Stays signed into your apps | Yes, which is what makes the bots possible | Not as it stands today |
| Setup effort | Easiest of the three he compares | No setup, but no persistence either |
| Trust barrier | You sign into your accounts on a machine you do not own | Familiar, and asks for less |
| His daily driver today | Not yet | Yes, with Codex |
The competitive read
His last strategic point is about the shape of the market rather than the product:
"It's just great to be in a world where Cursor and xAI are just as viable a competitor as OpenAI and Anthropic. I think Cursor may even have the edge if it can continue to support multiple models from all providers." (22:05)
The reasoning behind the edge is worth extracting: a product that is a harness rather than a model can route to whichever model is currently best, which is a structurally different bet from a lab shipping the interface to its own weights. If the harness is the product, model leadership becomes a supply question rather than an existential one.
He closes with the practicalities. Grok Bot is free to download, and he recommends trying it to get a glimpse of where this is going. He is putting the prompts from the video into the pinned comment. And an exclusive interview with the team on how they built Grok Bot is coming in the next few weeks.
"I'm really impressed by Grok Bot. I think the team really cooked here, and I can't wait to hear the story behind how they built this." (22:35)
How to reproduce all of this
The video is a tutorial, so here is the whole method compressed, in his order, with nothing added.
- Install Grok Bot and build the Advisor first. Tell it about your work and your life in a paragraph, then ask for five bots that would save you time or money, with a few sentences each.
- Have the Advisor create the working bots. Merge and rename its own suggestions in plain language rather than starting each bot from a blank prompt.
- Connect the plugins the bots need before you need them. Gmail, Google Drive, and anything with an MCP server (he uses Mercury for banking). Expect to sign in on the cloud computer, not on your laptop.
- Kick each bot off with a specific initial prompt, including the sources it should look at and the shape of the output.
- Expect the first output to be unusable. Iterate. Specify the report format explicitly, section by section. Constrain the time window if the domain is topical.
- Force a numbered list, and cap the length (he uses a maximum of ten items per category). Then approve or reject work by number.
- For anything destructive, put the guardrail in the prompt: do not move, delete, unsubscribe or cancel anything without my approval.
- Schedule the bot as a daily, weekly or monthly routine once the output is right.
- Push the output to where you already look. He routes reports to email so he never has to open the app.
- Give the bot a personality if it helps you read it. Marie Kondo is funnier and therefore more likely to be read than the same list from a robot.
Key takeaways
- The product difference is not the model, it is a persistent computer in the cloud with its own browser and OS that stays logged into your apps. Every automation in the video is downstream of that one property.
- Build an Advisor bot first and pin it to the top. One bot that knows your work and life can propose and then create all the others.
- Never one shot it. The reliable loop is: initial prompt, iterate on the output format, constrain the length, then schedule it. He runs this loop for every single bot.
- Ask for numbered lists. It turns approval into one short sentence and makes review possible at all.
- Write the guardrail into the prompt for any bot that can delete, unsubscribe or cancel: nothing happens without explicit approval.
- Agents beat alerts when the goal, not the query, is what you hand over. The concierge found a Tokyo round trip about $2,700 cheaper than the planned open jaw itinerary precisely because it was allowed to question the route.
- Marie Kondo did roughly 30 to 60 minutes of manual cleanup in about 5 minutes, including cancelling two live subscriptions, with a human approving each action.
- Push output to the channel you already open. He mails the X Scout report to himself rather than opening another app.
- The cloud computer is real enough to install games on, which is the clearest demonstration of what the agent can actually do with a machine, even though input lag makes Doom unplayable.
- Trust is the adoption bottleneck, not capability. Signing into your bank and your email on a machine you do not own feels different, even when the encryption story is correct.
- The price gap decides it for now: $200 a month against $20, so ChatGPT and Codex stay his daily drivers despite the cleaner UI.
Chapters
- 0:00 How Grok Bot is different from ChatGPT and Hermes
- 3:30 Advisor: Create and orchestrate other bots
- 4:30 YouTube Researcher: Find outlier videos and content ideas
- 7:00 X Scout: Surface the best posts from your feed and bookmarks
- 10:30 Marie Kondo: Clean your email and save money on paid subscriptions
- 15:00 Personal Concierge: Get alerts for better flight prices for your trips
- 17:30 Gamer: Install and play retro games
- 20:30 My honest take on privacy, pricing, and whether Grok Bot replaces ChatGPT
Notable quotes
"You know the meme of people walking around with their laptop half open to keep their agents running? You don't have to do that anymore if you're using Grok Bot." (1:21)
"Overall, I think this UX feels more like talking to a coworker than getting lost in a hundred chat threads." (3:10)
"What many people don't realize is that you can actually get one bot, in this case our Advisor bot, to help us create the rest of our bots." (4:14)
"By the way, this is how you should be working with AI to refine its output, instead of trying to one shot something." (5:52)
"Even though I have an unhealthy addiction to X, I still miss plenty of great content, or bookmark tweets that I never look at again." (7:12)
"Crucially, I told it to not move, delete, unsubscribe or cancel anything without my approval. You always want to review what it plans to delete before letting it remove anything." (11:12)
"It's always good to ask AI to give you its response in a numbered list, to make it super easy to just tell it to do things by referring to the number instead of typing everything over again. I use this pattern all the time." (12:40)
"It did all this in around five minutes, when it would have taken me probably 30 minutes to an hour to do all this manually." (14:00)
"Equip Foods no longer sparks joy. We release the prime protein subscription with gratitude." (14:35, Marie Kondo bot, in character)
"A simple Google price alert would not have found this." (16:35, on the $2,700 cheaper Tokyo round trip)
"Grok Bot is not replacing your gaming PC or GeForce NOW yet. But the fact that the agent can install and launch these games on its own computer gives you a sense of how open ended this could become." (19:10)
"How do I know that nobody else is seeing this screen on the virtual cloud computer?" (20:30)
"I think the AI agent platform that figures out trust will be the first to get mass adoption." (21:20)
"We're moving away from manually using our keyboard and mouse to do work on our laptops, to using our voice to orchestrate a bunch of agents that live in a dedicated cloud computer." (21:30)
"It's not quite my daily driver yet, because I think ChatGPT still offers more for $20 a month, while Grok Bot requires paying $200 a month to use on a regular basis." (21:45)
Resources mentioned
The product under test
- Grok Bot and xAI, the persistent cloud computer and the bots built on it. Free to download; he cites $200 a month for regular use.
What he compares it against
- ChatGPT, his current daily driver at $20 a month
- OpenAI Codex, the other half of his current workflow
- Cursor, which he argues may have the edge if it keeps supporting models from every provider
- OpenAI and Anthropic, the incumbents in his competitive read
- Hermes, the always on local agent he runs on a Mac Mini at home (he mentions it in passing and does not link it)
Apps and services the bots connect to
- Gmail, via the Gmail plugin, for the audit and for delivering reports
- Google Drive, for the abandoned and oversized file cleanup
- Google Flights, the price source for the concierge bot
- X, the source for the X Scout bot (follows, engagements and bookmarks)
- Mercury, his business bank, connected over MCP for subscription data
- Lovable and Equip Foods, the two subscriptions cancelled on camera
- YouTube, for the channel monitoring and comment mining
Games installed on the cloud computer
- Doom, installed and launched, unplayable because of mouse look
- Command & Conquer: Red Alert, installed
- Commander Keen, playable with input lag (Control is jump)
- GeForce NOW, the cloud gaming benchmark he says it is not replacing
- Windows 3.1, which the cloud computer's file manager reminds him of
People and namesakes
- Peter Yang, the creator, and his newsletter and podcast work at Behind the Craft
- Jenny Wen, on the design team he credits for the interface, previously interviewed on this channel
- Lee, who sent him the DM to test Grok Bot; Ben, who posted the most loved internal use cases; Riley, whose comparison post the X Scout surfaced
- Marie Kondo, the namesake of the cleanup bot
- A promised follow up: an exclusive interview with the team on how Grok Bot was built, coming to the channel in the following weeks
Where it stands
Everything above is his experience of a very new product, filmed as an early tester who was invited in by the team building it, and it is worth reading with that context rather than instead of it.
What is demonstrated on camera is solid. The bots exist, the reports arrive, the unsubscribes and cancellations actually execute, the flight comparison actually returns a number. He also shows the failures: three bad drafts before a usable report, a games experiment that half works, a cancellation flow that required him to type credentials twice.
What is claimed rather than demonstrated is the durability. A daily job that has run for a few mornings is not the same as one that survives a month of expired sessions, changed login flows and rate limits, and every automation here depends on staying signed in to services that have commercial reasons to log agents out. The $2,700 saving is a quoted comparison, not a booked ticket; he has not bought the flights yet.
The trust question he raises is the one that outlives the product. His own framing is the right one: he signed in because he is an early adopter, and he expects normal users to hesitate. Nothing in the privacy note he reads out is unusual for a cloud service, and the discomfort is not really about the encryption. It is that a browser session on a remote machine has no analogue in how most people think about their own computers, so there is no intuition to lean on when deciding whether to type a password into it.
And the price is the honest verdict, stated by the creator of the tutorial at the end of the tutorial: $200 a month against $20 is a wide gap, and he did not switch.


