IA & Agentes
Copilot Studio - Novidades [5] - Qual modelo escolher: Claude Sonnet 5, GPT-5.5 e as tags Deep, Auto e General
Copilot Studio - What's New [5] - Which model to pick: Claude Sonnet 5, GPT-5.5 and the Deep, Auto and General tags
Fala dataholics, último post da série de novidades. Uma pergunta pra começar: você sabe qual modelo está rodando no seu agente do Copilot Studio agora? Se a resposta é "sei lá, o que veio", então esse post vale seu tempo, porque a lista mudou bastante e a escolha mexe direto em latência, custo e qualidade de resposta.

O que veremos nesse post:
Onde troca o modelo e quantos lugares fazem isso
As tags General, Auto e Deep, que é o que realmente importa
Quem entrou, quem saiu e o que está experimental
A coluna Brasil, cross-geo e o que o admin precisa liberar
Onde fica a chave
Vai na página Overview do agente e na seção Model tem um dropdown. Você troca ali e pode alternar entre modelo de produção e experimental quando quiser.

Esse dropdown define o modelo primário, que é quem faz a orquestração generativa do agente, ou seja quem decide o que chamar e como responder. Só que existem outras três chaves separadas, cada uma com sua configuração: deep reasoning, generative responses e o prompt builder. Isso significa que você pode misturar, deixando o agente orquestrando num modelo rápido e um prompt específico rodando num modelo mais caro só quando precisa. Na prática é aqui que mora a economia.
As tags são o mapa
Antes de decorar nome de modelo, entende as três categorias, porque elas resumem o trade-off:
General - conversa do dia a dia, resumo, tradução, resposta FAQ com grounding e automação simples. Menor latência, menor custo, raciocínio raso a moderado.
Auto - roteia a pergunta dinamicamente, bom pra agente de helpdesk e de funcionário, onde a complexidade é imprevisível. Latência e custo variáveis, raciocínio adaptativo por turno.
Deep - raciocínio multi-passo com muita tool, análise de contrato e de política, troubleshooting que atravessa vários sistemas, síntese de documento longo com citação. Latência mais alta e custo mais alto.
Se o seu agente responde dúvida de RH, ele não precisa de Deep, e colocar Deep ali só vai deixar o usuário esperando e a fatura maior. Já um agente que lê política e cruza três sistemas antes de decidir sofre no General.
Quem está na lista hoje
O GPT-4.1 é o modelo Default na tabela de disponibilidade pública, ou seja é o que seu agente usa se você nunca mexeu. Quem está em GCC, GCC High ou DoD segue com o GPT-4o como default, então se você atende governo essa linha é diferente pra você. Em GA você tem GPT-5 Chat, GPT-5.5 Chat, Claude Sonnet 4.6 e Claude Sonnet 5 como General, mais Claude Opus 4.6 e Claude Opus 4.7 como Deep. Em preview aparecem GPT-5 Reasoning (Deep) e GPT-5 Auto (Auto).
DETALHE IMPORTANTE: o Claude Sonnet 5 só funciona em agente da nova experiência, aquela que eu abri no post [1]. Se você está na clássica, ele não vai aparecer pra você.
E teve saída também: GPT-4o e Claude Sonnet 4.5 estão como Retired. Modelo aposentado ainda dá pra usar por até um mês depois da aposentadoria, ligando a opção de continuar usando modelo retirado, e depois disso o agente cai no default. Se você tem agente antigo com modelo fixado, essa é a hora de revisar antes de tomar susto.
Na gaveta dos experimentais tem GPT-5.3 Chat, GPT-5.4 Reasoning, GPT-5.5 Reasoning, Mistral Medium 3.5 e Grok 4.1 Fast. Aliás, sobre o Grok, vale ler o aviso da própria Microsoft: a avaliação de segurança e IA responsável achou ele menos alinhado que os outros, com risco maior de gerar conteúdo prejudicial e nota pior em benchmark de jailbreak, incluindo categoria de dano que o sistema de content safety da Microsoft pode não cobrir. Achei corajoso publicar isso na doc do próprio produto, e é o tipo de informação que eu não esperava encontrar num catálogo de modelo. Fica o aviso: experimental não é pra produção, ponto.
Brasil, cross-geo e o que o admin precisa liberar
A tabela de disponibilidade tem coluna por região e o Brasil está lá. O ponto que interessa pra quem se preocupa com LGPD é a marca cross-geo: modelo com essa tag pode processar e armazenar dado fora da fronteira geográfica da sua organização. Vários modelos chegam ao Brasil justamente como GA cross-geo, incluindo o Claude Sonnet 5 e o GPT-5.5 Chat. Ou seja, disponível não quer dizer que o dado fica aqui.
E não é só o maker que decide. O admin controla três coisas:
Se modelo preview e experimental pode ser usado no environment.
Se a movimentação de dado entre regiões está ligada, que é requisito pra usar experimental.
Se modelo externo é liberado, e aí tem dois passos: habilitar no Power Platform admin center e liberar cada provedor no Microsoft 365 admin center, um por um, Anthropic, Mistral e xAI.
Reginaldo, e se eu publicar em preview só pra testar com uns usuários, sai de graça?
Não sai. A documentação é clara: se você publicar um agente com modelo experimental ou preview e as pessoas usarem, o uso é cobrado nas taxas normais. Teste é teste, mas a fatura chega igual.
Como eu escolheria
Meu jeito prático, e aqui é opinião: comece no default, meça com Evaluation antes de trocar qualquer coisa, e só depois suba de tag. Eu escrevi dois posts inteiros mostrando como um agente saiu de 40% pra 90% em Evaluation e a maior parte daquele ganho veio de instrução e knowledge, não de modelo. Trocar pra Deep sem medir é a forma mais rápida de dobrar custo e continuar com o mesmo problema. Se quiser ver aquele exercício, está no post de 40% a 90%.
Uma nota de quem está de olho no que vem: o GPT-5.6 já apareceu no Microsoft 365 Copilot em julho, no Word, Excel, PowerPoint, Copilot Chat e Cowork. Na lista do Copilot Studio ele ainda não está, então é questão de tempo até rolar pra cá também.
RESUMO
Modelo primário fica em Overview > Model, e existem chaves separadas pra deep reasoning, generative responses e prompt builder.
Escolha pela tag: General pra volume, Auto pra intenção imprevisível, Deep pra raciocínio pesado.
GPT-4.1 é o default. GPT-4o e Claude Sonnet 4.5 estão retired com um mês de tolerância.
Claude Sonnet 5 só na nova experiência.
Cross-geo significa dado podendo sair da sua região, e experimental publicado é cobrado normal.
Com esse post eu fecho a série de novidades. Ficou fora dela um monte de coisa que também saiu e que rende post próprio, tipo computer use em GA controlando browser e app desktop, o Windows 365 for Agents MCP server, voice agent integrando com Teams Phone Agent e o Entra Agent ID por agente. Comenta aí qual desses você quer ver primeiro que eu já boto na fila.
Referências:
Select a primary AI model for your agent
Print desse post: documentação oficial da Microsoft.
Fique bem e até a próxima.
#copilotstudio #claude #gpt5 #microsoft #ia #agentes #datainaction
Hey dataholics, last post in the what's new series. A question to start: do you know which model is running on your Copilot Studio agent right now? If the answer is "no idea, whatever came with it", then this post is worth your time, because the list changed a lot and the choice hits latency, cost and answer quality directly.

What we'll see in this post:
Where you switch the model and how many places do that
The General, Auto and Deep tags, which is what really matters
Who came in, who left and what's experimental
The Brazil column, cross-geo and what the admin has to allow
Where the switch lives
Go to the agent's Overview page and in the Model section there's a dropdown. You switch it there and you can move between production and experimental models whenever you want.

That dropdown sets the primary model, the one doing the agent's generative orchestration, meaning the one deciding what to call and how to answer. Except there are three other separate switches, each with its own setting: deep reasoning, generative responses and the prompt builder. Which means you can mix, leaving the agent orchestrating on a fast model and one specific prompt running on a pricier model only when it needs to. In practice that's where the savings live.
The tags are the map
Before memorizing model names, understand the three categories, because they sum up the trade-off:
General - everyday chat, summarizing, translating, grounded FAQ answers and simple automation. Lowest latency, lowest cost, shallow to moderate reasoning.
Auto - routes the question dynamically, good for helpdesk and employee agents where complexity is unpredictable. Variable latency and cost, reasoning adapts per turn.
Deep - multistep reasoning with lots of tools, contract and policy analysis, troubleshooting that crosses several systems, long document synthesis with citations. Higher latency and higher cost.
If your agent answers HR questions, it doesn't need Deep, and putting Deep there will only keep the user waiting and the bill bigger. On the other hand an agent that reads policy and crosses three systems before deciding suffers on General.
Who's on the list today
GPT-4.1 is the Default model on the public availability table, meaning it's what your agent uses if you never touched it. Anyone on GCC, GCC High or DoD still has GPT-4o as the default, so if you serve government that line reads differently for you. In GA you have GPT-5 Chat, GPT-5.5 Chat, Claude Sonnet 4.6 and Claude Sonnet 5 as General, plus Claude Opus 4.6 and Claude Opus 4.7 as Deep. In preview you get GPT-5 Reasoning (Deep) and GPT-5 Auto (Auto).
IMPORTANT: Claude Sonnet 5 only works in new experience agents, the one I covered in post [1]. If you're on classic, it won't show up for you.
There were exits too: GPT-4o and Claude Sonnet 4.5 are marked Retired. A retired model can still be used for up to a month after retirement, by turning on the option to keep using retired models, and after that the agent falls back to the default. If you have an old agent with a pinned model, now is the time to review it before you get a surprise.
In the experimental drawer there's GPT-5.3 Chat, GPT-5.4 Reasoning, GPT-5.5 Reasoning, Mistral Medium 3.5 and Grok 4.1 Fast. About Grok, by the way, it's worth reading Microsoft's own warning: their safety and responsible AI evaluation found it less aligned than the others, with higher risk of producing harmful content and worse scores on jailbreak benchmarks, including categories of harm that Microsoft's content safety system may not cover. I found it brave to publish that in their own product docs, and it's the kind of information I didn't expect to find in a model catalog. Consider it a warning: experimental isn't for production, period.
Brazil, cross-geo and what the admin has to allow
The availability table has a column per region and Brazil is there. The part that matters for anyone worried about LGPD is the cross-geo mark: a model with that tag may process and store data outside your organization's geographical boundary. Several models land in Brazil exactly as GA cross-geo, including Claude Sonnet 5 and GPT-5.5 Chat. So available doesn't mean the data stays here.
And it's not only the maker who decides. The admin controls three things:
Whether preview and experimental models can be used in the environment.
Whether data movement across regions is on, which is a requirement for experimental models.
Whether external models are allowed, and that takes two steps: enable it in the Power Platform admin center and allow each provider in the Microsoft 365 admin center, one by one, Anthropic, Mistral and xAI.
Reginaldo, and if I publish on preview just to test with a few users, is it free?
It isn't. The documentation is clear: if you publish an agent with an experimental or preview model and people use it, that usage is billed at the established rates. A test is a test, but the invoice shows up the same.
How I'd choose
My practical approach, and this is opinion: start on the default, measure with Evaluation before swapping anything, and only then move up a tag. I wrote two full posts showing how an agent went from 40% to 90% in Evaluation and most of that gain came from instructions and knowledge, not from the model. Switching to Deep without measuring is the fastest way to double your cost and keep the same problem. If you want to see that exercise, it's in the 40% to 90% post.
A note from someone watching what's coming: GPT-5.6 already showed up in Microsoft 365 Copilot in July, in Word, Excel, PowerPoint, Copilot Chat and Cowork. It isn't on the Copilot Studio list yet, so it's a matter of time until it rolls over here too.
RECAP
The primary model sits in Overview > Model, and there are separate switches for deep reasoning, generative responses and the prompt builder.
Choose by tag: General for volume, Auto for unpredictable intent, Deep for heavy reasoning.
GPT-4.1 is the default. GPT-4o and Claude Sonnet 4.5 are retired with a one month grace period.
Claude Sonnet 5 only in the new experience.
Cross-geo means data can leave your region, and a published experimental model is billed normally.
With this post I close the what's new series. A pile of other things that also shipped got left out and each one deserves its own post, like computer use in GA controlling browsers and desktop apps, the Windows 365 for Agents MCP server, voice agents integrating with Teams Phone Agent and the Entra Agent ID per agent. Tell me in the comments which one you want first and I'll put it in the queue.
References:
Select a primary AI model for your agent
Screenshot in this post: official Microsoft documentation.
Take care and see you next time.
#copilotstudio #claude #gpt5 #microsoft #ai #agents #datainaction
Gostou? Tem mais no YouTube e no LinkedIn.
Enjoyed it? There's more on YouTube and LinkedIn.