AI Agent News Today
Thursday, August 13, 2026SpaceXAI’s Grok Bot turns AI agents into persistent teammates for business apps
What changed: SpaceXAI opened early beta access to Grok Bot, a system of persistent AI agents where each bot runs on a dedicated cloud computer and can sign into existing applications and websites, even those without clean APIs or MCP endpoints. The agents can continue multi-step jobs after the user disconnects, coordinate with peer bots through shared context, and learn reusable workflows from a single demonstration, returning to the user only for approval or completion. Access is tied to premium subscriptions such as SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium on desktop and iOS, with business pricing positioned for heavy professional use.
Why it matters: This moves agents from “toy automations” to always-on teammates that can handle email, CRM updates, spreadsheets, and more across multiple tools without constant supervision. Founders and operators can begin shifting repetitive back-office tasks to autonomous agents, but must confront new issues around credential management, data access, and auditability across every app these bots log into.
Try/watch: Start with one tightly scoped workflow—such as inbox triage or CRM hygiene—and define clear guardrails for which accounts Grok Bot can access and what actions it may take, then monitor logs and approvals before expanding to more sensitive processes.
Nvidia ships Nemotron 3.5 Lightning and NeMo Switchyard for faster, smarter agent workflows
What changed: Nvidia released Nemotron 3.5 Lightning, an open 30-billion-parameter mixture-of-experts model with only 3 billion parameters active at any time, built on a hybrid Mamba-Transformer latent MoE design and tuned for high-volume specialized agent tasks. On the PinchBench agent benchmark, the model reportedly delivers up to 4× faster output token generation and around 30% faster agentic task completion than comparable models while matching accuracy on coding, research, and file-management workloads. Alongside it, Nvidia launched NeMo Switchyard, an open-source routing library that can dynamically choose between open, proprietary, and Nvidia models at each step of an agent workflow to optimize for quality, latency, or cost.
Why it matters: Builders no longer have to choose a single “one-size-fits-all” model for their agents; they can mix cheaper, faster models with heavier systems where quality matters, without hand-wiring every decision. This can cut serving costs and response times, making complex, multi-step agents more feasible for smaller companies and high-volume workflows.
Try/watch: Integrate Switchyard into a pilot agent that handles a full workflow—such as document analysis plus code changes—and benchmark latency and cloud costs against a single-model setup to see if dynamic routing pays off at your scale.
River AI raises $1.1B to power trainable personal agent stacks
What changed: River AI, founded by xAI co-founder Igor Babuschkin, closed a $1.1 billion round led by General Catalyst and AMP PBC, with strategic backing from Nvidia, AMD Ventures, Y Combinator, and Temasek. The company offers a training API that performs LoRA fine-tuning and reinforcement learning runs on frontier open-weight models, completing complex RL jobs in 15–20 minutes without requiring a dedicated infrastructure team. River claims its approach can deliver training at two-to-four-times lower cost than closed alternatives while keeping the models open-weight for downstream customization.
Why it matters: Personal and vertical agents will need continuous fine-tuning on proprietary workflows and feedback, and River is positioning itself as a “training backend” that lets teams iterate without building full ML infrastructure. Founders can potentially own their agent stack on open models while still achieving rapid RL-driven improvements in performance and behavior.
Try/watch: Identify one high-value workflow—such as sales follow-up or support triage—and design a feedback loop that could feed into River-style RL training, then compare the economics versus relying solely on closed, fixed-weights APIs.
Cloud.ru launches Agents Space and GigaAgent for everyday autonomous assistants
What changed: Cloud.ru introduced Agents Space, a dedicated environment for using and creating personal AI agents, anchored by GigaAgent, described as Russia’s first autonomous universal AI agent for everyday tasks. GigaAgent can manage calendars, work with documents, handle correspondence, research information, and generate reports and presentations, and it is built on the open-source Ouroboros self-developing AI agent project. Users can either choose from ready-made agents tailored to specific tasks or design their own, with new customers receiving a 4,000-ruble credit to experiment with Agents Space in both work and personal contexts.
Why it matters: This is a concrete example of a cloud provider turning “agentic AI” into a mainstream consumer and SMB service, rather than leaving it as a developer-only concept. Localized platforms like Agents Space can accelerate adoption by bundling agents, tooling, and credits, while also setting norms around how autonomous assistants should behave in everyday productivity work.
Try/watch: Treat Agents Space as a sandbox to map your daily routines—calendar, reporting, emails—into agent-managed workflows, and pay attention to how well GigaAgent handles multi-step tasks without micromanagement.
Consumer and compliant agents: DeepSeek V4 Pro and Specificity’s permission-based voice AI
What changed: DeepSeek released the formal API version of DeepSeek V4 Pro (DeepSeek-V4-Pro-0813), enhancing its agent capabilities and adding support for Responses API and Codex integration, with performance tests showing the new build approaching Fable 5 on multiple benchmarks. Chinese coverage notes that DeepSeek plans to raise pricing across its API portfolio soon, encouraging current users to plan consumption ahead of the hike. In mobile, Honor’s Robot Phone YOYO Pro mode uses a 300B+ on-device model to understand and break down long, casual spoken instructions and then autonomously execute cross-app workflows—such as ordering a cake, booking transport, and reserving a karaoke room in one request—while also driving more playful motion and camera behaviors. In parallel, Specificity announced a new permission-based architecture for its agentic AI Speed-to-Lead voice technology, giving site visitors tiered options that constrain what voice agents may do at low and medium levels (scheduling only) and expand to full Q&A and product discussion at high permission levels.
Why it matters: DeepSeek and Honor show agentic behavior moving directly into consumer apps and smartphones, where long, real-world tasks can be handed off in natural language and executed across multiple services with minimal user input. Specificity’s tiered permission model offers a blueprint for how marketers and sales teams can deploy aggressive voice agents while still honoring consent and TCPA-style telemarketing rules by binding capabilities to explicit user choices.
Try/watch: If you build consumer or marketing agents, study Specificity’s tiered permission structure and consider adopting a similar, transparent capability ladder so users know exactly what they’re authorizing. Mobile and app teams should experiment with longer, multi-step spoken commands and evaluate whether agentic execution can reduce friction in booking, shopping, or support flows.
Stop reading agent demos. Give one a job you repeat every week.
Describe the work, test the first result, and keep the agent available without running your own server.
Plans start at $29/month. Cancel anytime.
Hosted agent
OpenClaw or Hermes