Project MonetRequest demo
Home/Blog/GPT-6 Astra API & Pricing Guide

AI · Project Monet Briefing

GPT-6 Astra API & Pricing Guide

A practical guide to the GPT-6 Astra model ID, token pricing, long-context surcharge, reasoning controls, agent features and cost management.

Published 2026-09-04 · Updated 2026-09-04 · By Project Monet Editorial Team

GPT-6 Astra API pricing and agent-cost workflow guide

01

GPT-6 Astra API model ID

Use gpt-6-astra in API requests. OpenAI recommends the Responses API for reasoning and tool-heavy workloads, while the model page also lists Chat Completions support.

The model supports low, medium, high, xhigh and max reasoning effort. Astra does not support a none reasoning setting.

02

Current GPT-6 Astra pricing

At publication time, Standard text pricing per 1 million tokens is $10 input, $1 cached input, $12.50 cache writes and $50 output. Tool-specific features such as search or computer use can have additional per-tool charges.

Batch and Flex are priced at 50% of Standard rates. Fast mode is priced at 2x applicable rates where supported, and OpenAI says Fast mode is unavailable for Astra with EU data residency.

03

Long-context pricing and context limits

Astra has a 1,050,000-token context window and a 128,000-token maximum output. That can accommodate very large repositories and document sets, but using the full window is not automatically economical.

If a prompt exceeds 272K input tokens, OpenAI prices the full request at 2x input and cache rates and 1.5x output rates. Retrieval, prompt caching, compaction and tighter context selection can therefore materially affect cost.

04

Async tool calling and mid-turn steering

Async tool calling lets the model continue other work while your application runs a slow function or custom tool marked asynchronous. Your application still executes the tool and returns the result using the original call ID.

Mid-turn steering lets a user change requirements while a response is already running over WebSocket. Astra can preserve completed work and continue with the updated instruction instead of discarding the whole task.

05

Reasoning controls and cost management

Astra supports configuration updates that can raise or lower reasoning effort during a conversation while preserving the original cached prompt prefix. Higher effort can improve difficult planning or reasoning, but it can also increase latency and billed output-token usage.

For production systems, compare success rate, latency and full task cost rather than defaulting every request to max reasoning. Use cheaper models for routine work, prompt caching for stable repeated context, and Batch or Flex where latency requirements allow.

06

Rate limits and deployment checks

The current model page lists no Free-tier support. Published paid-tier limits range from Tier 1 at 500 requests per minute and 500,000 tokens per minute to Tier 5 at 15,000 requests per minute and 40,000,000 tokens per minute.

Sources

Primary and supporting sources

Facts were rechecked against the linked sources immediately before publication. Pricing, product availability and rollout status can change.

Project Monet

Useful signals. Clear decisions. Better digital work.

Project Monet turns relevant shifts in AI, creator tools and the web into practical context—and builds focused websites for businesses ready to grow.

Request a free homepage concept