Gemini Gets One Step Closer to Controlling Other Apps on Your Android Phone
Google's Gemini AI assistant appears poised to take a significant leap forward in functionality, moving beyond conversational responses to actively controlling applications on Android devices. Recent discoveries in beta software suggest the AI could soon perform practical tasks like ordering food or booking transportation through automated interactions with supported apps.
Evidence of 'Screen Automation' Capabilities
Code strings found in the Google app beta version 17.4.66 reveal references to a feature described as "Get tasks done with Gemini." The functionality appears to involve what Google may call "screen automation" - essentially allowing Gemini to view, scroll, and tap within applications to complete specific actions on behalf of the user.
The feature description within the beta software explains: "Gemini can help with tasks, like placing orders or booking rides, using screen automation on certain apps on your device." This suggests the AI would be able to launch applications like Uber or food delivery services and navigate through the ordering process autonomously once given a command.
How the Feature Would Work
Based on the discovered code and descriptions, the functionality would operate through a simple voice or text command. For example, a user could tell Gemini to "order me a ride to the office," and the AI would then open the Uber app, input the destination, select the ride type, and complete the booking process. Similarly, for food delivery, Gemini could navigate through supported apps to place orders from specified restaurants.
Google appears to be implementing safeguards and limitations for this functionality. The beta code includes warnings that "Gemini can make mistakes" and that users are "responsible for what it does on your behalf, so supervise it closely." Users would reportedly have the ability to stop the AI agent at any point during task execution and take over manually. Initially, support may be limited to a select number of applications rather than all apps on a device.
The Broader Context of Agentic AI
This development represents the natural progression toward what the industry calls "agentic AI" - artificial intelligence that doesn't just respond to queries but takes actions to accomplish goals. Google previously demonstrated similar capabilities through its Project Astra initiative, which showed Gemini's potential to control device interfaces and perform multi-step tasks.
The move toward agentic functionality aligns with broader industry trends where AI assistants are evolving from passive responders to active agents capable of executing complex workflows. This represents a significant shift in how users might interact with their devices, potentially reducing the need for manual app navigation for routine tasks.
Conclusion
The evidence from recent beta software strongly suggests Google is actively developing screen automation capabilities for Gemini that would allow the AI to control applications and perform practical tasks on Android devices. While initial implementation may be limited to specific apps and include user supervision requirements, this represents a meaningful step toward more autonomous AI assistance. As these features potentially roll out through experimental Labs programs, they could fundamentally change how users interact with their smartphones for everyday tasks like transportation and food ordering.
