What Grok Bot launched—and what it has not proved yet

Grok Bot is available in beta for SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium subscribers. The public product page lists Cursor Ultra at $200 a month for individuals and Cursor Premium Teams at $120 per seat a month. Enterprise access is still a waitlist.

Each account gets a persistent cloud computer with a browser, filesystem and terminal. Bots can use structured connectors when available or work through a site's ordinary interface. They can continue in the background, run scheduled routines and pass work to other Bots in a shared conversation.

SpaceXAI gives internal examples: updating CRM notes, drafting follow-ups, processing invoices, preparing demo environments and reproducing software bugs. Those examples show the intended scope. They do not establish a success rate. The launch includes employee and early-user testimonials, but no public benchmark, outside task set, failure count or comparison between completed work and human cleanup.

VentureBeat reached the same evidence boundary in its launch coverage: SpaceXAI did not release performance benchmarks for Grok Bot. For a beta that can change records inside logged-in software, a buyer should treat 'works inside the app' as a capability description. Reliability still has to be measured on the buyer's own job.

Watching the happy path does not teach the exception policy

The teach-by-demonstration feature records visible computer interaction for up to ten minutes and does not record microphone audio. After the demonstration, Grok Bot creates a skill that can be reviewed, edited, tested and eventually attached to a routine.

A demonstration can capture the route through a clean invoice: open the email, copy the vendor, enter the amount, attach the PDF and save the record. It does not automatically explain what to do with a credit memo, a duplicate invoice, a new bank account, a missing purchase order, a different currency or a total that disagrees with the attachment. Those cases may never appear during the ten minutes the Bot is watching.

SpaceXAI's documentation says this plainly: the learned skill is a draft. Users are told to add decision rules, failure handling and approval boundaries that may not be obvious from one example, then test the skill on a safe case before putting it on a schedule.

That is the right mental model. Showing the clicks teaches a first route. Writing the edge cases teaches the job.

Several Bots can have different names and the same keys

The product often speaks about a team of Bots, but the account boundary matters more than the roster. SpaceXAI's docs say every Bot on an account uses the same cloud computer. Browser cookies, signed-in sessions, files and command-line credentials are shared. Each Bot has a separate screen, but those screens are explicitly not separate security boundaries.

If an Expense Manager signs in to an accounting system, another Bot on the same account may be able to use that session. If one Bot saves a sensitive spreadsheet in the shared workspace, the other Bots can see it. Deleting a Bot removes its profile, conversation and routines; it does not necessarily remove the shared files or browser sessions it used.

This does not make the product unusable. It changes the setup question. A separate Bot is a way to organize responsibility, not a way to isolate access. If two jobs should not share a login or file, giving them different names is insufficient.

Grok Bot also requires cloud data storage and does not support Cursor's Legacy Privacy Mode. SpaceXAI directs users to manage privacy and training choices through the applicable Cursor account settings. Anyone considering private client, employee or financial data should settle that account-level data arrangement before the first demonstration, not after the routine has learned the task.

A practical first test for an AI assistant that learns by watching

Choose one recurring task with a clear finish and cheap mistakes. Drafting a weekly report from approved sources is safer than sending customer email. Reconciling a copied set of non-sensitive records is safer than changing live payment details.

Before the demonstration, write down five cases the normal example will not show. Include one missing source, one stale source, one duplicate, one value outside the usual range and one action that must stop for approval. This short list prevents the recording from becoming the entire specification by accident.

Demonstrate the clean case, then inspect the generated skill. Add the source of truth, the expected output, the five exception rules, what counts as partial completion and which actions require a person. SpaceXAI recommends keeping sending, purchasing, deleting, publishing and production changes behind approval.

Run two safe tests: one ordinary case and one exception. A Grok Bot test run performs real work—it can navigate websites, change files and call connected tools—so use copies or a staging account and keep write actions gated. Confirm that the Bot selected current inputs, stopped at the right point and made its failure visible instead of guessing.

Only then schedule it. For the first week, count accepted results, wrong actions, missed exceptions, questions, cleanup minutes and how often the person reopened the original material. The useful result is a smaller chore. A routine that saves ten minutes and creates fifteen minutes of checking is still unfinished.

Jun wants the shared login visible. Mara wants the refusal cases written down.

Jun Vega would put the shared-computer warning at the exact moment a person signs in: every Bot on this account can use this browser session. Separate screens can look like separate rooms, and most people will not read a security page before connecting an inbox. The interface should show the actual access boundary while the choice is being made.

Mara Vale is less worried about whether the Bot memorized the clicks than whether it learned when those clicks should stop. One successful demonstration contains no declined payment, angry customer, stale spreadsheet or suspicious attachment. Before the routine runs alone, she wants at least three written refusal or handoff cases beside it.

Jun is dealing with the access people can see. Mara is dealing with the exceptions a clean demonstration hides. Both are ordinary setup work, and both are easier to do before an always-on routine starts accumulating trust.

Teach the boring version first

The watch-and-learn idea is useful because many small teams have jobs they can demonstrate more easily than they can diagram. A ten-minute recording can turn tacit clicks into a draft someone can inspect. That is a better starting point than a blank automation canvas for a lot of people.

Keep the first job narrow, draft-only and easy to check. Make the source current, write down the odd cases and show the Bot where the human decision begins. Expand access after it handles those boring edges without quietly inventing a rule.

One clean demonstration is the beginning of training. The proof arrives later, when the strange invoice shows up and the routine knows enough to stop.