In this blog post Global Regional and Data Zone Deployments in Microsoft Foundry we will explain why a small deployment setting can create major consequences for compliance, availability and cost.
Many organisations choose an AI model, deploy it in Azure and assume the selected resource location controls where every prompt is processed. That is not always the case. The deployment type determines whether model processing can occur globally, within a defined data zone such as Asia Pacific, or only in one Azure region.
At a high level, Microsoft Foundry is a platform for building and managing business AI applications. It brings together AI models, agents, testing tools, security controls and connections to company data within the Microsoft Azure environment.
When an employee asks a Foundry application a question, the request is sent to a deployed AI model. The deployment type tells Microsoft how widely it can route that request across its Azure infrastructure.
What the deployment choice actually controls
Your deployment decision affects three important areas.
- Processing location โ where prompts and model responses can be handled.
- Capacity and availability โ how much Microsoft infrastructure can be used to serve requests.
- Cost and performance โ whether you pay for actual usage, reserve capacity or process work in batches.
It helps to separate data processing from data storage. Microsoft states that data stored at rest remains within the designated Azure geography. The deployment type primarily changes where live model processing, also called inference, can occur. Inference simply means the AI model reading a prompt and producing a response.
This distinction is important for Australian organisations. Creating a Foundry resource in an Australian Azure region does not automatically mean every Global deployment will process prompts only in Australia.
For a deeper look at that issue, read our guide to how Microsoft Foundry keeps AI data in the Asia Pacific region.
Global deployments prioritise scale and model access
A Global deployment allows Microsoft to process requests in any Azure region where the selected model is available. This gives the platform a larger pool of infrastructure from which to serve your workload.
For many businesses, that flexibility is useful. It can provide access to higher quotas, broader model availability and more capacity during periods of heavy demand.
Consider an internal assistant that helps employees draft generic emails, summarise public reports and prepare meeting agendas. If it does not handle sensitive client or regulated information, Global Standard may offer a sensible balance of cost, capacity and simplicity.
The trade-off is the processing boundary. Prompts and responses may be processed outside Australia and outside the Asia Pacific region. That may conflict with customer contracts, internal policy or a risk assessment covering sensitive business information.
Global does not automatically mean insecure. It means the organisation is accepting a wider geographic processing boundary in exchange for greater infrastructure flexibility.
Data Zone deployments provide a practical middle ground
A Data Zone deployment allows Microsoft to route requests across multiple regions, but only within a Microsoft-defined zone. Current zones include the United States, European Union and Asia Pacific.
For an Australian business using an Asia Pacific Data Zone deployment, processing can occur within the wider Asia Pacific boundary rather than being restricted to one Australian Azure region.
This option can be valuable when the business wants stronger geographic control but also needs more capacity and resilience than a single-region deployment may provide.
For example, a 200-person professional services firm might use an AI assistant to search policies, prepare first drafts and answer questions from an approved knowledge base. The firm may decide that global processing is too broad, while strict single-region processing unnecessarily limits model choice or capacity.
An Asia Pacific Data Zone deployment can provide a workable compromise. The organisation gains access to infrastructure across the zone while keeping processing inside a defined geographic boundary.
However, โAsia Pacificโ does not mean โAustralia only.โ Technology leaders should make sure this difference is understood by legal, compliance and risk teams before approval.
Regional deployments provide the narrowest boundary
A Regional deployment processes requests in the Azure region associated with the deployment. This provides the most tightly controlled processing location of the three options.
Regional deployment may be appropriate for workloads involving sensitive legal documents, health information, financial records, government data or contracts that specify a particular processing location.
The business cost is reduced flexibility. Not every model or model version is available in every Azure region, and regional capacity can be more constrained. A project may therefore face lower throughput, fewer model choices or a longer wait for capacity.
There is also a resilience question. Keeping processing in one region creates a clearer boundary, but it can increase dependency on that location. If the application is business-critical, review our practical guide to high availability and disaster recovery for Microsoft Foundry agents.
Location is only half of the deployment decision
Microsoft Foundry also asks how you want to consume capacity. This is separate from the global, data zone or regional processing boundary.
- Standard uses pay-per-token pricing. A token is a small unit of text processed by the model. This option is usually suitable for pilots, variable demand and workloads where usage is difficult to predict.
- Provisioned reserves processing capacity. It can provide more predictable throughput and less variation in response times for high-volume production applications.
- Batch processes large, non-urgent jobs asynchronously rather than responding immediately. It can suit tasks such as classifying thousands of documents overnight.
Microsoft currently offers combinations such as Global Standard, Data Zone Standard, Global Provisioned, Data Zone Provisioned, Regional Provisioned and several batch options. Model support and regional availability vary, so these combinations must be confirmed before the architecture is approved.
The naming can be confusing because โStandardโ may describe a consumption option while a regional pay-per-use deployment can also appear as โStandard.โ Focus on two separate questions: where can processing happen, and how will capacity be purchased?
A practical way to choose the right deployment
Start with the information being processed, not with the newest model.
- Classify the data. Decide whether the application will handle public, internal, confidential, personal or regulated information.
- Define the allowed processing boundary. Confirm whether processing can occur globally, within Asia Pacific or only in a nominated Azure region.
- Check the actual model. Confirm that the required model and version support your preferred deployment type.
- Estimate demand. Consider the number of users, expected request volume and periods of peak activity.
- Plan for failure. Decide what employees should experience if the selected model, capacity pool or region is unavailable.
- Record the decision. Document who approved the processing boundary, model, connected data sources and fallback arrangements.
Do not assess the model in isolation. A Foundry application may also connect to Azure Storage, Azure AI Search, databases, logging platforms and third-party services. Each component can have its own location, security and retention settings.
Access controls are equally important. Geographic boundaries will not protect sensitive information if employees, applications or external services have unnecessary permissions. Our guide to securing Microsoft Foundry with role-based access control and managed identity explains how to limit access and remove stored credentials.
Avoid making one choice for every AI project
A common mistake is forcing every workload into the same deployment pattern. This can either increase risk or create unnecessary cost and complexity.
A public-content marketing assistant may be suitable for Global Standard. An internal knowledge assistant could require an Asia Pacific Data Zone. A highly sensitive case-management application may need regional processing with reserved capacity.
The right answer depends on the data, business impact and contractual obligations of each workload. It should not depend on whichever option a developer happened to select during an early test.
Make the processing boundary a business decision
Global deployments generally provide the greatest flexibility. Data Zone deployments offer a balance between geographic control and access to wider capacity. Regional deployments provide the narrowest processing boundary but may limit model availability and resilience options.
CloudProInc helps organisations make these decisions with practical input from security, operations, compliance and finance. As a Melbourne-based Microsoft Partner with more than 20 years of enterprise IT experience, we focus on building Foundry environments that are supportable after the pilot is finished.
If you are not sure where your current Microsoft Foundry workloads process data, or whether the selected deployment type matches your risk requirements, we are happy to take a look โ no strings attached.
Discover more from CPI Consulting
Subscribe to get the latest posts sent to your email.