Grok 4.7
Grok 4.7 is xAI's reasoning chat model. It reads text and images and always
thinks before answering; you choose how hard with reasoning_effort. Call it
with the OpenAI SDK or with the native FlyMy.AI API, on the same wallet as
every other model.
| Endpoint id | flymyai/grok-4_7 |
| Model page | https://app.flymy.ai/models/flymyai/grok-4_7 |
OpenAI base_url | https://api.flymy.ai/api/v1/flymyai/grok-4_7/openai/v1 |
| Native route | POST https://api.flymy.ai/api/v1/flymyai/grok-4_7/predict |
OpenAI-compatible API
from openai import OpenAI
client = OpenAI(
base_url="https://api.flymy.ai/api/v1/flymyai/grok-4_7/openai/v1",
api_key="fly-***",
)
response = client.chat.completions.create(
model="grok-4.7",
messages=[
{"role": "system", "content": "Answer briefly."},
{"role": "user", "content": "Why is the sky blue?"},
],
reasoning_effort="low",
max_tokens=400,
)
print(response.choices[0].message.content)
print(response.usage)
The model is selected by the base_url; model is echoed back. The key is your
FlyMy.AI key as a Bearer token.
With stream=True, the answer arrives as one content chunk when Grok has
finished, followed by a chunk with finish_reason, the usage chunk and
data: [DONE]. Keep-alive comments flow while it thinks, so long answers do
not time out.
usage follows the OpenAI convention: completion_tokens includes the
reasoning tokens, itemized in completion_tokens_details.reasoning_tokens, and
prompt_tokens_details.cached_tokens shows the part of the prompt read from
the cache.
Native API
The native route takes form fields or a JSON body and answers with one Server-Sent Event carrying the result:
curl -X POST https://api.flymy.ai/api/v1/flymyai/grok-4_7/predict \
-H "X-API-KEY: fly-***" -H "Content-Type: application/json" \
-d '{"prompt": "Write a haiku about rain.", "reasoning_effort": "low"}'
output_data holds output (the answer), finish_reason, model,
stream_details (input_tokens, output_tokens, cached_tokens,
reasoning_tokens) and charge_usd, the price of the call.
Grok 4.7 does not use /predict/stream/. For long jobs, /predict/async/
works as for other models, except in private mode.
Parameters
| Parameter | Values |
|---|---|
prompt | One user message |
messages | Chat history in the Chat Completions format (a JSON array in form data); used instead of prompt. Text and image_url content parts |
system_prompt | Instructions for the whole chat, added unless messages already has a system message |
reasoning_effort | low, medium, high (default) or xhigh; none and minimal mean low |
max_tokens | Most tokens to generate, 1 to 128,000 (max_completion_tokens works too) |
temperature | 0 to 2 |
top_p | 0 to 1 |
private_mode | true keeps the prompt and the answer out of FlyMyAI storage and logs |
tools, tool_choice and response_format are not supported yet: the
OpenAI-compatible route rejects them with 400 unsupported_parameter.
Private mode
Send private_mode: true - extra_body={"private_mode": True} in the OpenAI
SDK, or a form or JSON field on the native route. FlyMy.AI then keeps neither
the prompt nor the answer: the request history shows the call, its token
counts and its price, without the text. Private mode works with synchronous
calls only; an async request with it is refused, because an async request is
replayed from storage.
xAI's own retention policy for API requests still applies.
From an AI assistant
An assistant connected to FlyMy.AI quotes Grok 4.7 with get_pricing, which
also returns its OpenAI base_url, and asks it with run_model
(endpoint_id: flymyai/grok-4_7, the same input as /predict). The result is
Grok's answer as text, followed by its token counts and charge_usd; clients
that read structured tool output find the answer in answer. private_mode
works here too, since the call is synchronous.
Pricing
| Prompt size | Input | Cached input | Output |
|---|---|---|---|
| Under 200K tokens | $2.00 / 1M | $0.50 / 1M | $6.00 / 1M |
| 200K tokens or more | $4.00 / 1M | $1.00 / 1M | $12.00 / 1M |
Reasoning tokens are billed as output. Each call is charged what xAI reports for
it, so cached prompt tokens are billed at the cached rate. The price of every
call is in charge_usd (native) and in your usage.
Errors
| What you see | Meaning |
|---|---|
422 on the OpenAI route | A parameter is out of range, or messages is malformed; the message says which |
400 unsupported_parameter | tools, tool_choice or response_format |
HTTP 200 with status_code: 400 from /predict/stream/ | Use /predict or the OpenAI-compatible route |
private mode does not support async requests | Send the private call synchronously |
403 Insufficient funds | Top up your balance |