Back to News
EnterpriseAI Understanding briefing

FDA says governed Databricks platform now supports more than 6,000 users

The FDA says its HALO data platform migrated thousands of users and pipelines to AWS GovCloud, creating a governed foundation for regulatory AI workloads.

By 5 min readRead the primary source
Source-provided image accompanying FDA says governed Databricks platform now supports more than 6,000 users
The short version

The FDA says its HALO data platform migrated thousands of users and pipelines to AWS GovCloud, creating a governed foundation for regulatory AI workloads.

What happened

The FDA built HALO, an enterprise data platform using Databricks on AWS GovCloud and Unity Catalog. The agency says the migration covered more than 5,000 users, 8,000 jobs and pipelines, and more than 4,000 notebooks without downtime.

Databricks published the account on September 1, 2026, describing the FDA’s presentation at the company’s Data + AI Summit. According to the source, the FDA built HALO, or Harmonized AI and Lifecycle Operations for Data, as a secure, governed foundation for analytics and artificial-intelligence work. The platform runs on Databricks in AWS GovCloud and uses Unity Catalog as a common governance layer for data and AI assets. Databricks presents the architecture as suitable for sensitive workloads subject to requirements including FedRAMP High and DoD IL5, while the FDA describes the system as a way to modernize data operations without interrupting regulatory work.

The FDA’s modernization began with a legacy-environment project in 2020, followed by a move to Databricks on AWS GovCloud in 2025. The source identifies three milestones: sponsorship for FedRAMP High authorization, the GovCloud migration and adoption of Unity Catalog. The agency says it migrated more than 5,000 users and more than 8,000 jobs and pipelines with zero downtime. It also refactored more than 1,000 data pipelines and more than 4,000 notebooks. The platform reportedly brought eight centers and 30 programs onto one environment and consolidated more than 40 data sources from application and submission systems.

Databricks says the FDA now has more than 6,000 users on the platform, up from roughly 500 in 2020, and expects that figure to exceed 10,000 by 2028. The FDA reports that SQL warehouse response times for business-intelligence workloads improved by more than 30%, compute costs fell by more than 20%, time spent on provisioning, permissioning and data sharing declined by more than 75%, and operational overhead fell by more than 35%. The platform also supports model serving and Genie, which the source says are being used to modernize legacy dashboards. HALO is integrated with Elsa, the FDA’s enterprise AI platform, and supports MARS, an initiative for analyzing structured and unstructured material in drug and device submissions. The source says AI use cases are in production and pilot stages, with more than 10 additional efforts in flight.

Source details: databricks.com

Why it matters

The deployment illustrates how a federal regulator is attempting to make AI usable within a high-security, multi-tenant data environment while preserving access controls, auditability and human review. The operational results are significant, but they are self-reported in a Databricks corporate blog and are not independently verified in the source.

The FDA’s account is relevant because it places AI deployment inside the operational constraints of a major federal regulator rather than presenting AI as a standalone demonstration. The agency handles scientific and regulatory work across many centers, offices, laboratories and product categories. The source argues that fragmented data, duplicated work and inconsistent pipelines had limited the ability of analysts and scientists to find and reconcile information. A shared, governed platform could reduce those barriers by making approved data easier to discover and share while preserving separate controls for different organizational units.

The multi-tenant design is central to the deployment. The source compares it with an apartment complex: centers share an underlying platform but retain separate secured spaces, permissions and policies. Unity Catalog is described as the layer that allows data sharing across centers while maintaining governance, lineage and auditability. That arrangement could be practically important for public-sector AI because a model or analytical tool may need information from multiple programs, while sensitive records still require restrictions on access and use. The source also highlights infrastructure-as-code, private networking, customer-managed keys and a compliance security profile as components of the broader security architecture.

The strongest evidence in the source concerns reported operational changes, not the quality or safety of AI outputs. Databricks attributes the performance and cost figures to the FDA, but the blog does not provide measurement periods, baselines, methodology or an independent assessment. It also does not report whether MARS has shortened review timelines, improved error rates or changed regulatory outcomes. The source says AI augments scientists and reviewers rather than replacing them, but it gives no details about approval gates, escalation procedures, audit findings or incidents. Those omissions do not invalidate the deployment, but they limit what can be concluded about its real-world effectiveness and safety.

What to watch next

Key unknowns include how the FDA measured its reported efficiency gains, which AI use cases are in production, and whether MARS improves review quality or speed in practice. The agency’s expected expansion beyond 10,000 users by 2028 will test whether its governance and human-review approach scales.

The first priority is verification of the reported results. Further FDA documentation or a public presentation could clarify the time periods behind the claims of faster queries, lower compute costs and reduced provisioning work, as well as how zero downtime was defined during the migration. It would also be useful to know whether the figures apply to the full platform or selected workloads. The source gives no absolute cost figures, service-level results or detailed account of outages, failed migrations or remediation work.

The next test will be the transition from infrastructure readiness to measurable regulatory use. The source says MARS can help reviewers analyze clinical-trial results, safety data, labeling and other information associated with drug and device applications, but it does not specify how broadly the system is used or which tasks remain experimental. The FDA’s stated goal of exceeding 10,000 users by 2028 will make access controls, lineage, permissioning and human review more difficult to manage across a larger user base. Future reporting should distinguish deployments in production from pilots and planned projects.

Security and governance claims also warrant continued scrutiny as the system expands. Databricks describes private connectivity, customer-managed encryption keys, compliance controls and Unity Catalog as parts of the architecture, but the source does not provide an independent security assessment or details about how permissions are tested. It does not say which models power the AI applications, how their outputs are evaluated, or how sensitive submission data is retained and monitored. The practical question is whether the FDA can preserve traceability and accountable human decision-making while making AI tools broadly available across centers and programs.

Related guides & quizzes

AI EthicsAI Models ExplainedFuture of AITest what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?