返回新聞
產品展示AI Understanding 簡報

GitHub Copilot 通过计算机使用预览添加桌面应用程序自动化

GitHub 宣布,Copilot 現在可以在公共預覽版中控制 macOS 和 Windows 桌面應用程序,讓開發人員描述所需的結果,並讓 AI 執行點擊、打字、捲動和其他 GUI 操作。

4 min readRead the primary source
Source-provided image accompanying GitHub Copilot adds desktop‑app automation via computer‑use preview
主要來源文件來源記錄
出版商
github.blog
來源連結
github.bloghttps://github.blog/changelog/2026-10-01-github-copilot-can-now-interact-with-desktop-apps
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

特點
模型用來進行預測的輸入變數。
測試一下自己AI 代理測驗
Source video from github.blog · shown with attribution.

發生了什麼事

GitHub released a public preview of “computer use” for GitHub Copilot, available in the Copilot CLI and the Copilot desktop app for macOS and Windows. The lets the AI read accessible content from any application, click controls, type or edit text, press keys, scroll, drag, and navigate multi‑step workflows across apps that lack APIs or command‑line interfaces. The preview includes a demo where Copilot automates an expense‑report workflow in Safari. Users must approve each interaction, can set permanent allowances for trusted apps, and on macOS must grant Accessibility and Screen Recording permissions. Organization administrators can disable the capability via policy settings.

GitHub’s announcement, dated October 1 2026, describes the computer‑use capability as a public preview. It is integrated into the existing Copilot CLI and the Copilot desktop application for both macOS and Windows platforms.

The AI can perform a range of actions that mimic a human user: reading accessible UI elements, clicking buttons, entering or editing text, pressing keyboard shortcuts, scrolling, dragging items, and moving through multi‑step workflows. The is designed for apps that do not expose APIs, command‑line interfaces, or other programmatic hooks.

Interaction is gated by user consent. Copilot prompts for approval before taking control of an app, and users can review or reset permissions at any time. On macOS, the also guides users through the required Accessibility and Screen Recording permissions. Enterprise administrators can disable the feature via organization‑level settings.

The preview includes a visual example where Copilot navigates an expense‑report workflow in Safari, demonstrating end‑to‑end automation of a real‑world task.

來源詳情: github.blog ↗

為什麼這很重要

The addition expands Copilot from code‑centric assistance to broader desktop automation, enabling developers to script repetitive GUI tasks without writing custom bots or macros. This could accelerate onboarding for legacy tools, reduce context‑switching, and lower the barrier for non‑technical staff to automate routine workflows. Because the works on any app that exposes accessibility data, it opens a pathway for AI‑driven integration with proprietary or legacy software that otherwise cannot be programmatically accessed. At the same time, granting an AI control over the desktop raises security and privacy considerations; organizations will need to manage permissions, audit logs, and policy controls to prevent unintended actions or data leakage.

By moving beyond code suggestions to direct desktop manipulation, Copilot addresses a long‑standing productivity gap for developers who must interact with GUI‑only tools. This could reduce the need for custom scripting or third‑party automation platforms.

The capability democratizes automation for legacy software that cannot be instrumented through traditional APIs, potentially extending the useful life of older enterprise applications.

Security and privacy implications are significant. Granting an AI the ability to control the desktop requires robust permission models, audit trails, and clear organizational policies to prevent accidental data exposure or malicious use.

The preview’s limited availability—requiring explicit user approval and organization‑level enablement—suggests GitHub is testing both technical feasibility and governance frameworks before a broader rollout.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
互動式概念檢查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下來看什麼

Key signals to monitor include adoption rates among individual developers and enterprise teams, the evolution of permission‑management features, and any reported misuse or security incidents. GitHub’s roadmap for expanding computer‑use to more complex multi‑app orchestrations will indicate how quickly the capability matures. Regulators and corporate IT departments may issue guidelines on AI‑driven desktop control, especially in environments handling sensitive data. Finally, competitor responses—whether other IDE or AI‑assistant vendors introduce similar GUI‑automation features—will shape the broader market for AI‑augmented productivity tools.

User and enterprise adoption metrics released by GitHub in the coming weeks.

Enhancements to permission granularity, logging, and admin controls that address security concerns.

Reports of any unintended behavior, bugs, or security incidents linked to the computer‑use .

Competitive moves from other AI‑assistant providers that may introduce similar desktop‑automation capabilities.

Regulatory or industry guidance on AI‑driven desktop control, especially in sectors handling regulated data.

相關指引和測驗

人工智慧代理人工智慧模型解釋AI 的未來什麼是人工智慧?測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?