Kembali ke Berita
ProdukAI Understanding pengarahan

AWS menambahkan inferensi lintas Wilayah khusus India untuk OpenAI GPT-5.6 di Bedrock

Amazon Bedrock sekarang merutekan OpenAI GPT-5.6 Permintaan Terra dan Luna antara Wilayah AWS di Mumbai dan Hyderabad, memungkinkan pelanggan untuk menskalakan inferensi sambil tetap menjaga perintah dan keluaran di India, menurut AWS.

5 min readRead the primary source
Primary-source image accompanying AWS adds India-only cross-Region inference for OpenAI GPT-5.6 on Bedrock
Dokumen sumber utamaSumber direkam
Penerbit
aws.amazon.com
Tautan sumber
aws.amazon.comhttps://aws.amazon.com/blogs/machine-learning/introducing-india-cross-region-inference-for-openai-gpt-5-6-models-on-amazon-bedrock/
Jenis sumber
Dokumen primer — pengumuman resmi, makalah, pengarsipan, atau halaman pihak pertama yang kita baca langsung.
Juga dikutip

Cerita terakhir direvisi

KonteksPahami ini dalam 60 detik

Mulai di sini

Istilah-istilah penting

Kesimpulan
Fase runtime saat model terlatih menghasilkan prediksi atau keluaran.
API (Antarmuka Pemrograman Aplikasi)
Cara terstruktur bagi satu sistem perangkat lunak untuk mengirim permintaan dan menerima tanggapan dari sistem lain.
RAG (Generasi Augmented Pengambilan)
Sebuah metode yang mengambil pengetahuan eksternal dan memasukkannya ke dalam generasi pada waktu inferensi.
Uji diri Anda sendiriChatGPT & Kuis LLM

Apa yang berubah sejak publikasi

  1. Pertama kali diterbitkan
  2. This source materially advances the existing AWS announcement about India-only availability of OpenAI GPT-5.6 Terra and Luna on Bedrock by detailing India geographic cross-Region inference between Mumbai and Hyderabad, the relevant profile IDs, supported APIs, monitoring, data-retention exception, IAM requirements, and the distinction from global routing.

Apa yang terjadi

AWS says Amazon Bedrock now supports India geographic cross-Region for OpenAI GPT-5.6 Terra and Luna. Requests can be routed between the Mumbai and Hyderabad AWS Regions based on capacity, while remaining within India. Both models accept text and image input, produce text output, and offer a 1-million-token context window.

AWS announced that Amazon Bedrock supports OpenAI GPT-5.6 Terra and Luna through India geographic cross-Region profiles. The two profile IDs identified in the post are in.openai.gpt-5.6-terra and in.openai.gpt-5.6-luna. The source says customers can invoke the models from either the Asia Pacific (Mumbai) Region, ap-south-1, or the Asia Pacific (Hyderabad) Region, ap-south-2. Bedrock then selects the destination India Region according to available capacity.

The announcement describes cross-Region primarily as a capacity mechanism. Instead of tying an application to the capacity of one Region, the profile permits requests to use a broader pool of compute across the two listed India Regions. AWS says this is intended to help maintain throughput and consistent performance during traffic peaks. The source provides no benchmark data, service-level guarantee, latency comparison, or independent assessment showing how much performance improves.

The GPT-5.6 models described in the post accept text and image inputs and return text. Each has a 1-million-token context window, which AWS says can support long documents, large code bases, and workloads combining text and images in a single request. The models can be called through the OpenAI Responses API, OpenAI-compatible Chat Completions, or Amazon Bedrock’s Converse API. AWS also documents use through the Bedrock console playground and says existing OpenAI SDK clients can be pointed at the Bedrock endpoint.

The post includes operational details for deployment. Billing and quota consumption are tracked against the account in the source Region, while CloudWatch and CloudTrail entries are recorded there. AWS says the India profiles use the bedrock-runtime endpoint, where Bedrock features such as Guardrails, intelligent prompt routing, and cross-Region are available. Existing workloads using the bedrock-mantle endpoint remain supported, but customers must use bedrock-runtime to adopt the India geographic profiles.

Detail sumber: aws.amazon.com ↗

Mengapa itu penting

The change targets organizations that need India data residency while also needing access to pooled cloud capacity. It could simplify deployment for workloads involving long documents, large code bases, mixed text-and-image inputs, retrieval systems, and AI agents. The source does not provide independent performance measurements, pricing, or evidence of customer adoption.

The practical significance is concentrated in data residency and capacity management. AWS says requests using the India geographic profiles are routed only between Mumbai and Hyderabad, so prompts and generated outputs may move between those two Regions but remain within the country. For organizations whose policies or sectoral obligations require Indian processing, that offers a different deployment option from a global profile that may send requests to supported commercial AWS Regions worldwide.

The feature may be relevant to financial-services, healthcare, and public-sector organizations, which AWS explicitly identifies as examples of customers with local data-processing requirements. The source does not establish that any regulator, customer, or independent auditor has approved the arrangement. It also does not explain how an organization should determine whether its own legal, contractual, or security requirements are satisfied by regional processing alone.

The model capabilities described by AWS could make the profiles useful for applications that repeatedly submit large context windows, including document extraction, code analysis, retrieval-augmented generation, and agent workloads. Prompt caching is supported under the India profiles, and AWS says cached reads are billed at a 90 percent discount compared with uncached input tokens. That is an AWS pricing claim in the source, not an independently verified estimate of total application savings; actual costs will depend on usage, cache behavior, output volume, and other services.

Centralized monitoring may also reduce administrative work. AWS says invocation counts, token counts, latency, throttles, and errors are published per profile, while CloudTrail records an inferenceRegion field showing whether a request was processed in Mumbai or Hyderabad. This could help customers audit geographic routing without correlating logs across Regions. The announcement does not state whether these controls provide complete visibility into every downstream processing or support function outside the model-inference path.

Interactive Mechanism

Mekanisme Interaktif: Cara Kerja Sebenarnya

Jelajahi teknologi yang mendasari di balik perkembangan ini secara interaktif.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Pemeriksaan Konsep Interaktif+10 Points
ChatGPT & LLMs Quiz

What is a common training objective for an autoregressive language model?

Apa yang harus ditonton selanjutnya

Organizations will need to verify regional availability, pricing, quotas, latency, retention terms, and IAM configuration before production use. India-only profiles differ from Bedrock’s global profiles, which can route GPT-5.6 requests to AWS Regions worldwide. AWS also says content flagged by abuse-detection classifiers may be retained for offline detection despite its default zero-data-retention model.

Masalah pertama yang harus diverifikasi adalah ketersediaan dan kapasitas aktual. AWS mengarahkan pelanggan ke dokumentasi ketersediaan model regionalnya dan mengatakan bahwa perutean bergantung pada kapasitas, namun postingan tersebut tidak mempublikasikan kuota, latensi yang diharapkan, target waktu aktif, atau perbandingan dengan inferensi Wilayah tunggal. Tim produksi harus menguji beban kerja mereka sendiri, terutama permintaan multimodal yang besar dan lonjakan lalu lintas, sebelum berasumsi bahwa perutean lintas Wilayah akan memberikan tingkat kinerja tertentu.

Penanganan data memerlukan pembacaan pengecualian yang cermat. AWS mendeskripsikan Amazon Batuan Dasar menggunakan model tanpa retensi data secara default, artinya input dan output biasanya tidak disimpan. Namun, sumber tersebut mengatakan konten yang ditandai oleh pengklasifikasi deteksi penyalahgunaan otomatis untuk model tertentu, termasuk GPT-5.6, dipertahankan untuk deteksi penyalahgunaan offline. Pengumuman tersebut tidak menentukan periode penyimpanan, kriteria klasifikasi, proses penanganan, atau cakupan akses, sehingga detail tersebut masih belum diketahui untuk penerapan sensitif.

Konfigurasi keamanan dan akses mungkin menjadi kendala penerapan. AWS mengatakan peran IAM harus diberikan akses ke profil inferensi, model dasar di Wilayah sumber, dan model dasar di setiap Wilayah tujuan. Organisasi yang menggunakan Kebijakan Kontrol Layanan harus mengizinkan ap-south-1 dan ap-south-2; memblokir salah satunya dapat menyebabkan inferensi lintas Wilayah gagal. Tim juga perlu memeriksa apakah penyedia identitas, proses kredensial, pendekatan kunci API, dan kontrol jaringan yang ada mendukung konfigurasi yang terdokumentasi.

Terakhir, pelanggan harus membedakan profil geografis India dari profil global. AWS mengatakan profil global mendukung model GPT-5.6 termasuk Sol, Terra, dan Luna tetapi dapat mengarahkan permintaan ke seluruh dunia untuk kapasitas maksimum. Profil India adalah pilihan yang relevan untuk menjaga inferensi di dalam negeri. Sumber tidak menyatakan apakah semua fitur model, parameter, tingkat harga, atau varian masa depan akan tersedia secara identik di kedua mode perutean, sehingga diperlukan pemeriksaan dokumentasi yang berkelanjutan.

Panduan & kuis terkait

ChatGPT & LLMModel AI DijelaskanAgen AIUji pengetahuan Anda — coba kuis AI gratisCari istilah AI di glosarium kamiIkuti pelacak rilis model AI

Pembaruan dan koreksi

Kisah kanonik ini diperbarui ketika peristiwa yang berkembang berubah secara signifikan. URL dan tanggal publikasi aslinya tidak pernah berubah.

  • This source materially advances the existing AWS announcement about India-only availability of OpenAI GPT-5.6 Terra and Luna on Bedrock by detailing India geographic cross-Region inference between Mumbai and Hyderabad, the relevant profile IDs, supported APIs, monitoring, data-retention exception, IAM requirements, and the distinction from global routing.
Lihat log koreksi publik
Apakah ini berguna?