Skip to content
beginner

Make Your First Programmatic LLM API Request

The four-line API call is not the hard part. The hard part is the key, the bill, and the error message you cannot read. So let's get all three out of the…

Published 2026-10-03Updated 2026-10-0410 min read
Serene beach with golden sand and clear blue sky along a calm coastline.
Serene beach with golden sand and clear blue sky along a calm coastline. Photo by Matheus Bertelli on Pexels.

The four-line API call is not the hard part. The hard part is the key, the bill, and the error message you cannot read. So let's get all three out of the way in one sitting.

By the end of this tutorial, you will have a working Python script that sends one request to a hosted large language model, prints the generated text, and shows you the token counts behind it. You will also know what to do when the first run fails — because it often does, and the failure is usually more informative than the success.

We are deliberately not locking into one vendor. You will use a provider-neutral client so the same call shape works across providers. That shared shape is the abstraction. It is not a promise that switching providers is free — more on that boundary in a moment.

What You Are Actually Sending

A Python script sends a model name and message through a client to a hosted model. An environment-stored API key supplies authentication separately. The response returns generated text, a finish reason, and token usage.
The prompt is request data; the API key authenticates the request, and the response includes useful metadata alongside generated text.

Before any code, get the mental model right. An LLM API call is an ordinary authenticated HTTP request. It has an endpoint, headers, a JSON body, and a JSON response. If you have ever called any REST API, nothing about the transport will surprise you.

What is unusual is the payload. The response is generated, not looked up. That single fact is why prompt wording, token limits, latency, and cost become your problem instead of the provider's.

The request body carries two things you care about right now:

  • A model name — a string that selects which model handles the request.
  • A messages list — an ordered list of role/content pairs, typically a system message for instructions and a user message for the actual input.

The response carries the generated text plus metadata: why generation stopped, and how many tokens were consumed.

Provider-neutral clients exist because every major provider converged on roughly this shape. The client library translates your call into whatever the provider expects. You write one call; the model string picks the destination.

Note: Model identifiers, package versions, and response field names change over time. Treat every string in this tutorial as a placeholder to verify against your provider's current documentation, not as a permanent fact.

Set Up Python, a Client, and One Provider

You need three things: an isolated Python environment, a client library, and one provider account.

Isolate the environment

Assume Python 3.11 is installed. Create a virtual environment so this tutorial does not touch your system interpreter:

python3.11 -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate

Install a provider-neutral client

We will use LiteLLM as the worked example because it exposes a consistent interface across providers. Check its current documentation for the exact package name and version before installing:

pip install litellm

Get one API key

Pick one hosted provider. Create an account, then generate an API key from its dashboard. Follow that provider's current documentation for this step — key formats, prefixes, and dashboard locations change, and a copied snippet from an old blog post is the fastest way to start with a broken key.

Store the key in an environment variable

Never paste the key into your source file. An API key is a password. Anyone who has it can spend your credits.

export OPENAI_API_KEY="your-key-here"

The exact variable name depends on the provider. LiteLLM reads provider-specific variables, so check which one your chosen provider expects.

Before you commit anything, add your environment file to .gitignore:

echo ".env" >> .gitignore

Warning: Set a spend or usage limit in your provider dashboard now, before the first request. Not after the first surprise bill. This tutorial sends one short request, but a runaway loop in a later experiment will not announce itself.

To be clear about the boundary: this is one request, one provider, no retries, no streaming, no production concerns. That is the point.

Knowledge check

Check your understanding

Answer this question before you continue.

A script needs a provider key, and you want to avoid committing that secret. Which approach follows the article?
Scenario Interpretation

Focus: Configure provider credentials without embedding secrets in source code.

The Smallest Request That Works

Here is the entire script. Save it as first_call.py.

import os
from litellm import completion

response = completion(
    model="openai/gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain what an API is in one sentence."}],
)

print(response.choices[0].message.content)

Run it:

python first_call.py

Expected output is one sentence of generated text, something like:

An API is a defined interface that lets one piece of software request data or actions from another.

The exact wording will differ every run. That is the success criterion: you got generated text back, not a stack trace.

Now the lines, one at a time.

import os is there because you will read the key from the environment in a moment. The client picks it up automatically, but you will want os for the next section.

model="openai/gpt-4o-mini" selects the destination. The prefix before the slash tells the client which provider to route to. Change the prefix and the model name, and the same call hits a different provider — provided you have also configured that provider's credentials and the model identifier is one it currently supports. The call shape is portable; the setup is not. That distinction is the whole reason the abstraction is useful and the whole reason it can mislead you.

messages=[...] is the conversation. A single user message is the minimum. Add a system message when you want to set behavior:

messages=[
    {"role": "system", "content": "You are terse. Answer in one sentence."},
    {"role": "user", "content": "Explain what an API is."},
]

The key is read from the environment by the client library. You never pass it as a literal. If you find yourself typing api_key="sk-..." into a file, stop — that file is one git push away from being a public secret.

Knowledge check

Check your understanding

Answer this question before you continue.

You keep the call shape but change the model string to route to another provider. What else must be true for that request to work?
Comparison Reasoning

Focus: Identify what must be true for a provider-neutral call to work with a different provider.

Read the Response, Not Just the Text

The text is the headline. The rest of the response is the story.

Print the whole object once, structured, so you see its real shape instead of trusting a summary:

import json
print(json.dumps(response.model_dump(), indent=2, default=str))

You will see something like this, simplified:

{
  "choices": [
    {
      "message": {"role": "assistant", "content": "..."},
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 14,
    "completion_tokens": 22,
    "total_tokens": 36
  }
}

Three fields matter right now.

finish_reason tells you why generation stopped. "stop" means the model finished normally. "length" means it hit the output limit and the answer may be cut off mid-sentence. A truncated answer looks like a bad model but is almost always a max_tokens setting you did not set high enough.

usage.prompt_tokens and usage.completion_tokens are the quantities you need to estimate cost. Input tokens are what you sent; output tokens are what came back. They are not the bill itself — the bill is those counts multiplied by the selected model's current per-token pricing, which differs across models and can change. Log the counts from day one, then combine them with the model's pricing page to see what a call actually costs. Cost stops being abstract the moment you can see it per call.

Common mistake: Reading only response.choices[0].message.content and throwing the rest away. You lose the truncation signal and the cost signal in the same line of code.

Field names and nesting vary by provider and client version. Verify against current docs rather than assuming the shape above is universal.

Knowledge check

Check your understanding

Answer this question before you continue.

A response reports `finish_reason` as `length` and includes prompt and completion token counts. Which interpretation matches the article?
Scenario Interpretation

Focus: Interpret a response's finish reason and usage fields to recognize truncation and estimate cost.

When the First Run Fails

First-run errors are normal. They are also the fastest way to learn what the API actually does. Read the status code and the error body before changing code — the message is evidence, not noise.

SymptomLikely causeFix
Authentication failure (401)Key missing, misspelled, expired, or not exported into the shell running the scriptRe-export the variable in the same terminal, then rerun
Model not found / invalid request (400)Model identifier does not match what the provider currently offers, or request shape is wrongCopy the exact model string from current provider docs
Rate limit (429)Too many requests too fastSlow down; retry with backoff
Insufficient quota (402)No credits on the accountAdd credits or switch models
Timeout or provider error (408, 500–503)Provider-side issue; your request may have been fineRetry with backoff rather than rewriting the prompt
Context too large (400/413)Input exceeded the model's windowShorten the prompt or choose a larger-context model

Two of these deserve a closer look because beginners conflate them.

A rate limit means you are asking too fast. The fix is timing. Insufficient quota means you have no credits. The fix is money. They are different problems with different fixes, and the status code tells you which one you have.

A timeout is not a prompt problem. If the request timed out, the model may have been mid-generation. Retrying with backoff is correct. Rewriting your prompt because of a network hiccup is wasted work.

Tip: When something fails, print the status code and the raw error body. Most client libraries attach the provider's response to the exception. That body usually names the exact problem.

Knowledge check

Check your understanding

Answer this question before you continue.

A request fails with a 429 rate-limit error. According to the article, what should you try first?
Debugging

Focus: Choose a response to a rate-limit failure that differs from the response to insufficient quota.

One Small Change to Prove You Understand It

A copied script is not knowledge. Change one variable at a time and watch the consequence.

Change the prompt text. Replace the question with something personal to you. If the output changes, your input is genuinely reaching the model. If it does not, something upstream is wrong.

Set a low output limit. Add max_tokens=10 to the call and rerun. The goal is to try to induce truncation and watch finish_reason change from "stop" to "length". A short prompt may finish inside ten tokens, in which case the finish reason stays "stop" and nothing is broken — you just did not push hard enough. Ask for something longer, or drop the limit further, until you see "length". Truncation stops being mysterious the moment you trigger it on purpose.

Swap the model string. Point at a second model from the same provider and compare latency and usage counts. Same call shape, different destination — that is the value of a provider-neutral client, and the reminder that credentials and model support still have to line up.

Log the usage numbers. Append them to a file alongside the answer:

with open("usage.log", "a") as f:
    f.write(f"{response.usage.total_tokens} tokens\n")

Cost becomes observable instead of abstract. One variable changed, one observation recorded.

Where This Fits Next

The rule I would give any beginner here: get one request working end to end before adding anything else. Retries, caching, routing, and output validation are separate problems with their own failure modes. Layering them onto a request you have not yet seen succeed is how debugging turns into archaeology.

This tutorial covered one request, one provider, no production concerns. The natural next concerns are distinct topics, not extensions of this script:

  • Handling failures reliably when a request matters and the network does not cooperate.
  • Validating model output before another piece of code consumes it.
  • Choosing between models when quality, latency, and cost pull in different directions.

A first-run error is not a sign you did something wrong. It is the API telling you what it actually expects — which is more than any tutorial can. Run the script. Read the error. Fix the assumption. Then make the second request.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

You set `max_tokens=10`, but a short response still reports `finish_reason` as `stop`. What is the best interpretation?
Question 1 of 2Misconception Check

Focus: Interpret a low output-token limit without assuming every short response is truncated.

A beginner's first request has not succeeded yet. Which next step best follows the article's guidance?
Question 2 of 2Scenario Interpretation

Focus: Choose a safe sequence for learning an API request before adding more complex behavior.

Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.