What happened
OpenAI says Australian law firm Gilbert + Tobin has expanded its use of ChatGPT Enterprise and Codex beyond individual assistance into defined, multistep operational workflows. The firm says ChatGPT is used across operations, marketing, recruitment, finance, technology, business transformation and parts of legal practice, while Codex supports selected compliance checks, audit reporting, file preparation and internal technology work.
OpenAI’s September 1 case study says Gilbert + Tobin first introduced ChatGPT Enterprise to operations teams in a controlled setting, then expanded access to marketing, business development, recruitment, finance, technology, business transformation and parts of the legal practice. The firm says adoption was encouraged through visible leadership, role-specific demonstrations, internal videos, showcases and use-case stories rather than mandated usage targets. As of June 2026, OpenAI says 87% of enabled ChatGPT seats were active, more than twice the adoption rate the firm typically sees for other tools. The source does not define “active,” identify the number of enabled seats or provide a comparison methodology.
The firm says it established guidance covering approved tasks, permitted inputs and review of outputs before broadening use. OpenAI’s account also cites assessments of contractual protections, role-based access, data-processing requirements and administrative controls. Gilbert + Tobin says an OpenAI environment with Australian data residency increased its confidence in expanding access while meeting internal and client expectations. The source does not specify the retention period, audit-log design, model configuration, incident process or the categories of client information permitted in the system. Those omissions matter because the firm handles confidential legal and business information.
ChatGPT is described as supporting recruitment research and data extraction, reference processing, pitch preparation, industry reports, spreadsheet analysis, documentation, scripts, training materials and internal guidance. OpenAI says one recruitment workflow fell from approximately four hours to about 20 minutes, while reference processing saves an estimated 25 minutes per candidate. The firm produces 400 to 500 pitches each year, and says ChatGPT helps teams synthesize prior materials and tailor responses. Executives also use a custom GPT built from approved examples of the CEO’s writing, priorities and professional background to pressure-test ideas. The source says this tool does not make decisions or speak for the CEO.
Codex is presented as the next stage of the rollout, handling defined multistep work under human review. OpenAI says it prepared audit reports covering 300 entities, avoiding a full day of manual work across the workflow, and checked and renamed 1,100 files for upload into another system, replacing work that previously took days. The firm also used Codex to help build a Python application from requirements and an existing specification, and to create a monitoring “watchtower” for its AWS environment that consolidates operational signals, supports diagnosis and remediation actions, and escalates issues requiring human attention. For selected conflict, anti-money-laundering, politically exposed-person and know-your-customer checks, the firm says Codex completes research and processing before producing a report for human sign-off; selected checks that took three to eight hours now take minutes, with the source’s summary giving about five minutes.
Why it matters
The account offers a concrete example of generative AI moving into operational work at a professional-services firm that handles confidential information. It also illustrates the limits of the deployment: Gilbert + Tobin says employees remain responsible for constraining tasks, checking outputs and approving final work. The reported efficiency gains are company claims, and the source does not provide independent testing, error rates, costs or detailed audit results.
The deployment is significant because it concerns the operating layer of a high-trust professional service rather than a low-stakes demonstration. Recruitment research, pitch preparation, audit reporting, compliance screening and internal infrastructure monitoring are repetitive or information-heavy tasks that can consume substantial staff time. If the firm’s estimates are representative, the tools could change how support teams allocate time without directly replacing the legal judgment used to advise clients. The source, however, reports outcomes from one firm and does not establish that the same gains would apply elsewhere.
Gilbert + Tobin’s description also shows that adoption depends on organizational design, not only model capability. Senior leaders modeled use, teams received demonstrations tied to their roles, and the firm framed AI as a tool requiring employee judgment. The stated controls—approved tasks, restrictions on inputs, access management, data-processing review and human approval—are especially relevant in legal services, where an incorrect summary, missed conflict or unsupported compliance conclusion could create professional and commercial risks. The source claims these controls support confidence, but it does not show how they performed in practice or whether any failures have occurred.
The case illustrates a shift from asking a chatbot for an answer to assigning an AI system a bounded workflow with intermediate steps and a defined deliverable. That shift can increase productivity, but it also concentrates risk: an error may be repeated across hundreds of entities or files before a reviewer notices it. The reported time savings therefore should not be read as evidence of accuracy or legal sufficiency. The source provides no independent evaluation of Codex’s research, file handling, compliance screening, audit reports or AWS remediation suggestions, and it does not disclose how much human time remains in each process.
What to watch next
The firm plans to broaden access to ChatGPT Work and Codex after adding controls for wider use. The important questions are whether human review remains substantive as workflows expand, how often systems produce errors or require correction, and what records are retained for client, privacy and regulatory oversight. The source also leaves unclear how widely these tools are used in legal practice itself compared with administrative functions.
Gilbert + Tobin says it intends to extend access to ChatGPT Work and Codex once it has put controls in place for broader use. It describes a longer-term goal of an interconnected working environment in which employees can draw on approved organizational context, carry out operational tasks and deliver finished outputs without manually moving between systems. That would represent a larger integration surface than the workflows described here. It will be important to establish which systems are connected, what permissions AI receives and whether actions are reversible.
Future evidence should include performance measures beyond elapsed time. Useful indicators would be error and correction rates, false positives and false negatives in compliance checks, the share of outputs requiring substantial rewriting, escalation frequency, and the quality of audit trails. The source does not provide these figures, nor does it say whether the reported savings include review, exception handling, data preparation and system-maintenance time. Without those details, the efficiency claims remain directional rather than a complete assessment of value.
The expansion also raises practical questions about confidentiality, accountability and professional responsibility. Australian data residency may address one organizational requirement, but the source does not explain all applicable data flows, client-consent arrangements or contractual limits. As Codex moves from producing drafts to taking actions such as renaming files, diagnosing infrastructure issues or processing compliance information, the firm will need clear boundaries for authorization and human sign-off. The central test will be whether governance scales at least as quickly as the workflows and whether users can reconstruct what the AI did, what a person changed and who approved the result.