Skip to content

The OpenAI API Before LangChain: Roles, Temperature, Max Tokens and Streaming

Site Console Site Console
18 min read Updated Oct 11, 2026 AI & Tools 0 comments

Why Start Underneath

This post does not use LangChain. That is deliberate.

Every framework is a set of decisions made on your behalf, and the only way to evaluate those decisions is to know what they replaced. Learn ChatPromptTemplate before you have ever assembled a message array by hand and you will know how to use it but not why it exists — and when it behaves unexpectedly, you will have no model of the thing underneath to reason with.

So this is the baseline, in more detail than the rest of the series will need. Everything below was verified against openai 3.26.0 on Python 3.10 or newer, by reading the SDK source rather than its documentation.


First Steps

Two APIs, and Which One to Learn

OpenAI's Python SDK exposes two paths to a text completion.

client.chat.completions.create() is the older one. Most tutorials still use it because most tutorials predate the alternative. It still exists and still works.

client.responses.create() is the current one, and it is the one to learn. New capability lands there, and it is what LangChain's OpenAI integration now builds on — one of the breaking changes in post 2 was langchain-openai storing Responses API items in message content by default.

Everything below uses the Responses API.

Installing and Authenticating

pip install openai

The key goes in the environment, never in your source:

export OPENAI_API_KEY="sk-..."

For local development a .env file is more convenient, as long as it is in .gitignore:

# pip install python-dotenv
from dotenv import load_dotenv
load_dotenv()

Then the client finds it with no arguments:

from openai import OpenAI

client = OpenAI()   # reads OPENAI_API_KEY from the environment

Configuring the Client

The constructor takes more than most code uses, and three of its parameters are worth setting deliberately in anything that runs unattended:

client = OpenAI(
    timeout=30.0,       # seconds; default is the SDK's own, which is generous
    max_retries=3,      # default is 2
    organization=None,  # for accounts with multiple orgs
    project=None,       # for per-project usage attribution
)

max_retries defaults to 2, which means the SDK is already retrying failed requests for you with exponential backoff before you ever see an exception. Worth knowing when you are debugging latency: a request that took nine seconds may have been three attempts, not one slow call.

timeout matters because the default is long enough that a hung request can hold a worker for an uncomfortable stretch. For a user-facing path, set it to something you are willing to make someone wait.

There is also base_url, which points the client at a compatible endpoint — an Azure deployment, a local server, a proxy — and data_residency, which is the client-side switch for the regional processing that carries a price premium, as post 3 noted.

For a one-off override without building a second client, client.with_options(...) returns a copy with changed settings. client.with_raw_response and client.with_streaming_response give you the HTTP response alongside the parsed one, which is where rate-limit headers live.

The First Call

response = client.responses.create(
    model="gpt-5.6-luna",
    input="Explain retrieval in one sentence.",
)

print(response.output_text)

Three things in those five lines.

input accepts a bare string when you have nothing but a question. It also accepts a list of messages, which is the next section.

output_text is a convenience property. It walks the response's output items and concatenates the text ones — the real structure is richer, and you will care about that the moment you use tools, but not yet.

Model strings change with every release. Read them from the provider's model reference rather than copying them out of a blog post, including this one.

Reading the Response Properly

output_text is not the whole object. The fields you will actually use:

response.id              # 'resp_...' — needed for previous_response_id
response.model           # which model actually served it
response.status          # 'completed' | 'incomplete' | 'failed' | ...
response.output          # the structured output items
response.usage           # token counts
response.incomplete_details  # why it stopped, if it stopped early

status is the one most people never check, and it has six possible values: completed, failed, in_progress, cancelled, queued and incomplete. A response can arrive with text in it and a status that is not completed — more on that under max tokens.

The usage object maps exactly onto post 3's arithmetic:

response.usage.input_tokens
response.usage.input_tokens_details.cached_tokens     # prove your caching works
response.usage.output_tokens
response.usage.output_tokens_details.reasoning_tokens # where surprise bills hide
response.usage.total_tokens

Log those from day one, aggregated per feature. Without them, "our AI costs went up" is a mystery; with them, it is a line item.

Handling Failures

The SDK raises typed exceptions, and they want different responses:

import openai

try:
    response = client.responses.create(model="gpt-5.6-luna", input="Hello")
except openai.AuthenticationError:
    raise                                   # your key is wrong; retrying cannot help
except openai.BadRequestError as e:
    raise                                   # your request is malformed; fix the code
except openai.RateLimitError as e:
    ...                                     # back off and retry later
except openai.APITimeoutError:
    ...                                     # already retried max_retries times
except openai.APIConnectionError:
    ...                                     # network; retry with backoff
except openai.InternalServerError:
    ...                                     # their side; retry with backoff

The useful split is between errors that are your fault and errors that are transient. AuthenticationError, BadRequestError, PermissionDeniedError and NotFoundError will fail identically forever, so retrying them burns time and quota. RateLimitError, APITimeoutError, APIConnectionError and InternalServerError are worth retrying — and the client already retried them twice before you saw the exception.


System, User and Assistant Roles

A Conversation Is a List

Instead of a bare string, input takes a list of messages, each with a role:

response = client.responses.create(
    model="gpt-5.6-luna",
    input=[
        {"role": "developer", "content": "You are a terse assistant."},
        {"role": "user", "content": "What is a vector database?"},
    ],
)

Four roles are accepted on input: system, developer, user and assistant.

What Each Role Is For

user is what the person typed. Everything that comes from outside your application lives here, and that boundary matters: text in the user role is data the model should consider, not instructions it must obey. Putting untrusted input anywhere else is how prompt injection works.

assistant is what the model said previously. You send it back so the model can see its own history. You can also write it yourself — putting words in the assistant's mouth to demonstrate a pattern, which is what few-shot prompting does, and post 5 builds properly.

system and developer both carry instructions from you, the person building the application, rather than from the person using it. They are the channel for your rules. The distinction between the two names is a provider-level detail that has shifted over time; developer is the newer spelling for the same job, and both are accepted.

Two Places to Put Instructions

The Responses API also has a top-level instructions parameter:

response = client.responses.create(
    model="gpt-5.6-luna",
    instructions="You are a terse assistant. Answer in one sentence.",
    input="What is a vector database?",
)

Functionally this overlaps with a developer message. The practical difference is structural: instructions is a separate argument, so it stays out of the message list, which keeps your conversation history clean when you start appending turns.

The arrangement worth adopting is instructions for the constant behaviour contract, and the message list for the actual conversation. That is also the arrangement post 3 asked for on cost grounds — a stable prefix at the front is the part the provider can serve from cache at a tenth of the price.

Multi-Turn, and What It Costs

A model has no memory. It is a function from a list of messages to the next message, and it knows nothing about the previous call. Continuity is something you construct by sending the history back every time:

conversation = [
    {"role": "user", "content": "How do I reset my password?"},
    {"role": "assistant", "content": "Click 'Forgot password'."},
    {"role": "user", "content": "It didn't send the email."},
]

response = client.responses.create(
    model="gpt-5.6-luna",
    instructions="You are a support assistant.",
    input=conversation,
)

Now connect that to post 3. Every turn resends the whole conversation, so a chat that has run for thirty exchanges pays input tokens on all thirty, every time the user types. Cost does not grow linearly with conversation length; it grows with the square of it.

That single fact drives a lot of real design work — summarising older turns, dropping the middle, capping history length. The Responses API also exposes prompt_cache_key for influencing how prefixes are grouped for caching. Note that prompt_cache_retention is deprecated in favour of prompt_cache_options, which is the kind of detail you only catch by reading the SDK.

Letting the Server Hold the History

Pass previous_response_id and the provider threads the conversation for you:

first = client.responses.create(model="gpt-5.6-luna", input="My name is Dang.")

second = client.responses.create(
    model="gpt-5.6-luna",
    input="What is my name?",
    previous_response_id=first.id,
)

This is convenient and it moves your conversation state onto someone else's infrastructure. Which of those two sentences matters more depends entirely on your application — and on whether you are comfortable with store defaulting to retaining the exchange server-side.


Creating a Sarcastic Chatbot

A deliberately strong personality is the clearest possible demonstration of what a system prompt is: not a suggestion about tone, but a behaviour contract that governs every turn. We will build it in three passes, and it comes back in posts 5 and 7a rebuilt at higher levels of abstraction.

Pass One: The Persona

from openai import OpenAI

client = OpenAI()

PERSONA = (
    "You are a support assistant who is always technically correct and "
    "relentlessly sarcastic. Answer accurately first, then add exactly one "
    "dry remark."
)

response = client.responses.create(
    model="gpt-5.6-luna",
    instructions=PERSONA,
    input="How do I reset my password?",
)

print(response.output_text)

That works, and it is also where the first lesson lives.

Pass Two: The Guardrails

Run the version above a dozen times and the drift shows up. Sarcasm aimed at a frustrated user stops being funny, and "technically correct" occasionally becomes a refusal to answer at all. A persona is a direction, not a constraint, so the constraints have to be written down:

PERSONA = (
    "You are a support assistant who is always technically correct and "
    "relentlessly sarcastic.\n"
    "\n"
    "Rules:\n"
    "- Answer the question accurately and completely FIRST.\n"
    "- Then add exactly one dry remark. One, not three.\n"
    "- Never be cruel, and never mock the user for not knowing something.\n"
    "- Never refuse to help. Sarcasm is the delivery, not the content.\n"
    "- If you do not know, say so plainly, without the joke."
)

Every one of those lines is doing work. Remove "never be cruel" and the persona drifts somewhere you would not ship. Remove "answer accurately first" and you get a joke where an answer should be. Remove "exactly one" and the model, given permission to be funny, keeps going.

This is the real lesson of system prompts, and it is worth absorbing before the framework hides it: the same lever that gives you a voice gives you the failure mode that comes with it. The rules are not politeness boilerplate; they are the boundaries of the behaviour you actually asked for.

Pass Three: A Working Conversation

A complete, runnable script, with the history management and failure handling that the earlier passes skipped:

import openai
from openai import OpenAI

client = OpenAI(timeout=30.0, max_retries=3)
MODEL = "gpt-5.6-luna"
MAX_TURNS = 12          # cap history: cost grows with the square of length

history: list[dict] = []

def ask(question: str) -> str:
    history.append({"role": "user", "content": question})

    try:
        response = client.responses.create(
            model=MODEL,
            instructions=PERSONA,       # constant prefix: cacheable
            input=history,              # the varying part
            temperature=0.9,            # a persona wants some variety
            max_output_tokens=300,
        )
    except openai.RateLimitError:
        history.pop()                   # do not poison the history with a failed turn
        return "Rate limited. Try again shortly."

    if response.status == "incomplete":
        reason = response.incomplete_details.reason if response.incomplete_details else "unknown"
        history.pop()
        return f"Response was cut off ({reason}). Try a narrower question."

    answer = response.output_text
    history.append({"role": "assistant", "content": answer})

    del history[:-MAX_TURNS]            # keep the most recent turns only

    print(f"[{response.usage.input_tokens} in / {response.usage.output_tokens} out]")
    return answer

if __name__ == "__main__":
    while True:
        try:
            question = input("you> ")
        except (EOFError, KeyboardInterrupt):
            break
        if question.strip() in {"exit", "quit"}:
            break
        print("bot>", ask(question), "\n")

Four things in there are the difference between a snippet and something you could leave running.

The persona lives in instructions, not in the history, so it is a stable prefix that caching can reach and it never gets trimmed away.

A failed turn pops the user message back off the history. Otherwise a rate-limited question stays in the conversation forever, and the model sees a question that was never answered.

del history[:-MAX_TURNS] caps the growth. Without it, the quadratic cost from earlier is unbounded and the thirtieth question costs thirty times the first.

And the token counts print on every turn. That habit is the cheapest observability you will ever add.


Temperature, Max Tokens and Streaming

Temperature

A model does not choose the next token; it produces a probability distribution over all possible next tokens. Temperature reshapes that distribution before a token is sampled from it.

Low values sharpen it, so the most likely token almost always wins. High values flatten it, so less likely tokens get a real chance.

response = client.responses.create(
    model="gpt-5.6-luna",
    instructions="You name software products.",
    input="Suggest a name for a database migration tool.",
    temperature=1.2,
)

The accepted range is 0 to 2, and the practical mapping is simple. Anything with one correct answer — classification, extraction, routing, structured output — wants a low temperature. Anything where variety is the point wants a higher one. A sarcastic persona genuinely benefits from around 0.9, because the same joke every time stops being a joke.

Two caveats that trip people up.

Zero is not determinism. It makes output much more stable, not byte-identical across calls. Floating-point non-determinism and infrastructure variation remain, so do not build a cache key or a test assertion on the assumption.

Parameter support varies by model. Reasoning models in particular may ignore or reject sampling parameters. Check the reference for the model you are actually calling rather than assuming every parameter applies everywhere.

top_p, and Why Not Both

top_p is the alternative, called nucleus sampling. Instead of reshaping the whole distribution, it truncates it: keep the most likely tokens until their probabilities sum to top_p, then sample only from those.

# keep only the tokens making up the top 10% of probability mass
response = client.responses.create(model="gpt-5.6-luna", input="...", top_p=0.1)

Both parameters control the same thing — how adventurous the sampling is — by different mechanisms. The standard advice is to change one and leave the other at its default, because tuning both at once makes the effect of either impossible to reason about.

There is also top_logprobs, between 0 and 20, which returns the most likely alternatives at each position. It is a debugging tool rather than a generation setting, and it is genuinely useful when you want to know whether the model was confident or barely preferred its answer.

Max Output Tokens

max_output_tokens caps generation. Not the total, not the input — only what comes back.

response = client.responses.create(
    model="gpt-5.6-luna",
    input="Summarise the CAP theorem.",
    max_output_tokens=150,
)

It is a hard stop, not a hint. The model does not plan a shorter answer to fit; it is cut off mid-flow when it hits the ceiling, which can leave you with a truncated sentence or — worse, if you asked for JSON — invalid syntax.

So the parameter needs a partner: detect when it fired. This is the part almost no tutorial covers, and the SDK gives you exactly what you need.

if response.status == "incomplete":
    print(response.incomplete_details.reason)

status is one of completed, failed, in_progress, cancelled, queued or incomplete. And when it is incomplete, incomplete_details.reason is one of:

max_output_tokens   your ceiling stopped it
max_messages        the conversation exceeded a message limit
content_filter      the content was filtered
steered             generation was redirected

Checking status is the difference between handing a user half a sentence and telling them the answer did not fit. A response with status == "incomplete" still has text in it, and output_text will return that text happily — so if you only read output_text, truncation is silent.

Ask for brevity in the prompt when you want brevity. Use max_output_tokens as a safety rail against a runaway generation, and since output tokens cost roughly six times input tokens, treat it as your one hard guarantee about the cost ceiling of a single call.

There is also truncation, set to "auto" or "disabled", which governs what happens when the input is too long for the context window rather than the output.

Streaming

A model generates one token at a time, so the full answer takes as long as the full answer takes. Streaming does not make it faster — it makes the wait visible, which users read as responsive rather than broken.

stream = client.responses.create(
    model="gpt-5.6-luna",
    input="Write a two-sentence product description for a password manager.",
    stream=True,
)

for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)

The critical detail is that the Responses API streams typed events, not raw text fragments. You are not iterating over strings; you are iterating over an event sequence. The SDK defines 66 event classes. Text arrives as response.output_text.delta events carrying a delta string, interleaved with lifecycle events for the response and for each output item.

Filtering on event.type is not boilerplate you can skip — it is how you avoid printing the scaffolding.

The events worth handling in a first implementation:

response.created              the response exists; nothing generated yet
response.output_text.delta    a text fragment, in .delta
response.output_text.done     this text item is finished
response.completed            the whole response finished successfully
response.incomplete           it stopped early — check the reason
response.failed               it failed
response.refusal.delta        the model is declining, streamed the same way

A more complete loop looks like this:

chunks: list[str] = []

for event in stream:
    if event.type == "response.output_text.delta":
        chunks.append(event.delta)
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        usage = event.response.usage
        print(f"\n[{usage.input_tokens} in / {usage.output_tokens} out]")
    elif event.type == "response.incomplete":
        print(f"\n[cut off: {event.response.incomplete_details.reason}]")
    elif event.type == "response.failed":
        print(f"\n[failed: {event.response.error}]")

answer = "".join(chunks)

Note where the usage data is. While streaming, token counts are not available until the terminal event, because they are not known until generation stops. If you need to log cost per request and you stream, you have to accumulate it from response.completed rather than from the call's return value.

Streaming also changes your error handling in a way that is easy to miss. A failure halfway through a stream has already sent half an answer to the user, which a plain request-response call never does. Partial output is a state your application has to have an opinion about: discard it, keep it with a marker, or retry and replace it. There is no default that is right for every product.


🧭 What's Next

  • Post 5: Model Inputs in LangChain — the same sarcastic chatbot, rebuilt with typed message objects and prompt templates, plus few-shot examples that steer a model by showing rather than telling. This is where you see exactly what the first abstraction buys and what it costs.

Related

Leave a comment

Sign in to leave a comment.

Comments