Back to News
ProductAI Understanding briefing

Votee AI is building Cantonese models for Hong Kong’s public services and businesses

Fortune reports that Hong Kong startup Votee AI is retraining open-weight models on Cantonese data for banks, universities and government departments, aiming to address gaps in mainstream AI systems’ local and cultural knowledge.

By 6 min read
AI-generated editorial illustration accompanying Votee AI is building Cantonese models for Hong Kong’s public services and businesses
The short version

Fortune reports that Hong Kong startup Votee AI is retraining open-weight models on Cantonese data for banks, universities and government departments, aiming to address gaps in mainstream AI systems’ local and cultural knowledge.

What happened

Fortune reports that Hong Kong-based Votee AI is building Cantonese-focused language models by retraining open-weight models from developers including Meta and Alibaba. The company says it has expanded its Cantonese training corpus from about 100 million tokens to more than 500 million through online scraping, community and university contributions, previous data-business content and synthetic data. Its models are reported to have about 70 billion parameters, and the company estimates training costs of roughly $250,000.

Fortune reports that Votee AI is a Hong Kong startup focused on Cantonese and other languages that it describes as underserved by leading AI systems. According to the report, the company takes open-weight models from developers such as Meta and Alibaba and performs additional training with Cantonese data. Votee sells the resulting systems to banks, universities and government departments. The report presents the company’s work as a response to the dominance of English and Mandarin in the current AI market, rather than as an attempt to build a general frontier model that competes directly on overall scale.

Fortune reports that Cantonese differs from Mandarin in grammar and vocabulary, particularly in Hong Kong, where speakers may switch between Cantonese and English within the same sentence. The article says more than 80 million people speak Cantonese, a population comparable to the number of Korean speakers and larger than the number of Italian or Thai speakers. Yet the report says the supply of standardized written data, especially colloquial Cantonese, is comparatively limited. Fortune cites HKCanto-Eval, a benchmark developed by researchers at Kyushu University, the Education University of Hong Kong and the local AI community hon9kon9ize. The benchmark, which Fortune says was sponsored by Votee, found that mainstream models could handle everyday Cantonese reasonably but routinely failed on cultural and local knowledge.

The reported training process combines several sources. Fortune says Votee scrapes online material, including content from Radio Television Hong Kong, and also receives data from communities and universities. The company draws on material from its earlier big-data business and creates synthetic Cantonese datasets for training. According to Fortune, those efforts grew the corpus from 100 million tokens to more than 500 million. The startup’s models are described as having around 70 billion parameters. Chief executive Pak-Sun Ting told Fortune that Votee uses between 500 million and 1 billion tokens to train its models, compared with the trillions used for English-language models, and estimates the cost at roughly $250,000. Fortune does not independently verify those figures or describe the model’s architecture, release terms, evaluation methodology or availability to the public.

Read the primary source: fortune.com

Why it matters

Cantonese is spoken by more than 80 million people, but Fortune reports that colloquial and localized written data is relatively scarce. Mainstream models can handle everyday Cantonese but perform poorly on cultural and local knowledge, according to the HKCanto-Eval benchmark sponsored by Votee. Better local-language systems could affect public-facing uses in education, healthcare, policing and government services, while also supporting broader efforts to build AI that is less dependent on English and Mandarin.

Fortune’s reporting illustrates why language coverage is a practical AI issue, not only a matter of cultural preservation. Ting told the publication that Cantonese is used in education, healthcare and police communications, arguing that systems that do not handle those contexts are of limited use. If accurate, the company’s focus points to a gap between broad multilingual claims and the ability to understand local terminology, institutions, social conventions and mixed-language speech or writing. The consequences could be significant wherever an AI system summarizes information, supports staff or communicates with residents in Cantonese.

The report also places Votee’s work within the wider movement toward so-called sovereign AI. Fortune describes related efforts involving Indonesia’s Sahabat AI, AI Singapore’s SEA-LION project and South Korea’s state-backed effort to select domestic foundation-model providers. The common objective is greater control over local data, models and applications rather than complete dependence on overseas providers. Ting told Fortune that owning every part of the AI supply chain is difficult, but that governments could prioritize control of foundation models and applications. He also argued that many public-sector tasks may not require frontier-scale systems and could be handled by smaller local models, potentially with more powerful English- or Chinese-language models used behind the scenes.

That model has practical tradeoffs. Local systems may improve linguistic fit and give institutions more control over sensitive deployments, but a smaller or specialized model may have narrower capabilities and still depend on foreign model providers, chips or infrastructure. Fortune reports that Votee works with models from MiniMax and SenseTime and can use Nvidia chips, underscoring that local-language ownership does not necessarily mean technological independence across the full stack. The article also describes Votee as a for-profit company whose revenue currently exceeds its costs, according to Ting, with client contracts providing much of its funding. Fortune reports that Allan Zeman is an advisor and that Votee is in active discussions with AI Singapore, but it does not independently confirm the company’s finances, contracts, customer deployments or the status of those discussions.

What to watch next

The key unanswered question is whether Votee’s models perform reliably outside the company’s own claims and the benchmark it sponsored. Fortune does not provide independent evaluations, error rates, a public model release, pricing, customer numbers or details about data licensing and privacy safeguards. Further reporting should examine performance on real Cantonese tasks, the provenance and consent status of scraped material, deployment scale, and whether the company’s discussions with AI Singapore lead to a concrete partnership.

The first priority is independent performance testing. Fortune reports the HKCanto-Eval findings and Votee’s claims about Cantonese understanding and reasoning, but the article does not provide scores, comparisons across specific models, sample tasks, failure rates or evaluation by an unaffiliated auditor. Useful scrutiny would test colloquial Cantonese, code-switching with English, regional vocabulary, public-service terminology and culturally specific questions. It should also assess whether improvements persist when the system encounters unfamiliar topics rather than material resembling its training data.

Data governance is another unresolved issue. Fortune says Votee uses online scraping, RTHK content, community and university contributions, prior business data and synthetic data. The source does not establish which material was licensed, how contributors gave permission, whether personal information was removed, or how copyright and broadcaster rights are handled. Those questions matter especially for systems intended for healthcare, education, policing and government use. Future disclosures should clarify the corpus’s provenance, consent and retention policies, as well as safeguards against reproducing private or copyrighted material.

The company’s commercial and regional expansion will determine whether the project becomes a durable public capability or remains a specialized service business. Fortune reports that Votee sells to institutional customers, says its revenues exceed its costs, and hopes to expand from Hong Kong into Southeast Asia and eventually work on endangered languages in East Asia, North America and Africa. The source does not provide customer counts, deployment volumes, public access, pricing, service-level commitments or evidence of adoption. Watch for a public model release, independently verified contracts, results from real deployments, a confirmed AI Singapore partnership and evidence that the system can deliver reliable benefits without overstating what a 70-billion-parameter Cantonese model can do.

Related guides & quizzes

What is AI?ChatGPT & LLMsAI Models ExplainedAI EthicsTest what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?