Cogitation.ai
creAItive.llm · Python library + book

Build it.
Run it.
Make it better.

The experience of a machine learning team, now available to anyone.

creAItive.llm is a Python library for running and improving LLM applications. It brings reliability, cost tracking, user feedback, and quality evaluation into one interface.

By Geoff Hulten, author of Building Intelligent Systems.
Working draft · Looking for early readers and builders.

The four pieces that make it work.
  1. 01
    Separate the work from the model.

    Keep your application stable as models, providers, and execution choices change.

  2. 02
    Make operations visible to your agent.

    Execution records and CLI reports show how the application is running.

  3. 03
    Learn from how people use the app.

    Connect model results to the choices users already make, as well as feedback they give directly.

  4. 04
    Give improvement a repeatable process.

    Observe, evaluate a change, deploy it, and monitor the results.

Why this exists

You can build it now.
What happens next?

One person can now build things that used to require an entire machine-learning team. But the experience of operating those systems hasn’t spread nearly as far as the ability to build them.

creAItive.llm assumes you’re going to live with your application. Run it, maintain it, find where it fails, and make it better. It gives those practices a place in the API and a set of tools you and your coding agent can use from the beginning.

You decide what’s worth building and what a good result means. The library helps you preserve the evidence and establish the process for getting there over time.

How creAItive.llm makes that practical

Built for the life
of your application.

Four pieces that work together: a stable execution boundary, operational records, connections to user outcomes, and tools for deliberate improvement.

01 / Separate the work from the model

Keep the application stable.
Change how the work runs.

Your code asks for a support answer. A named policy selects the models, settings, timeouts, recovery, and fallback that perform that work. Those decisions can evolve in one place as the application grows.

Start with your coding agentExample prompt
You

I want to use creAItive.llm in this application. Start here:

https://cogitation.ai/creaitive/llm/

Check how to get access. Once we have the library, read its AGENTS.md and help me wire up one feature, with execution logging and a model policy I can change independently.

Then show me how we’ll track costs and learn whether the results help our users.

Currently available through early access. The package includes agent instructions, API documentation, and CLI tools.
02 / Make operations observable

Your agent can see
how the system is running.

With execution logging enabled, the runtime keeps track of the calls, attempts, models, costs, latency, and recovery. You and your agent get a consistent set of CLI reports to inspect that history.

Ask where the money is going, what’s slowing down, or whether fallback is keeping a feature working. The records are already organized to answer those questions.

Inspect the running applicationCLI
# Get a view of operational health.
python -m creAItive.llm report overview \
  --log-dir logs

# Break down the model spend.
python -m creAItive.llm report costs \
  --log-dir logs

# See which failures led to fallback.
python -m creAItive.llm report fallbacks \
  --log-dir logs
The same reports are available to you and your coding agent. Add --json for structured output.
See sample output Two weeks of real use

Scroll right for more

$ python -m creAItive.llm report overview --by requested_policy --period total \
  --log-dir logs \
  --start 2026-08-30T00:00:00-07:00 \
  --end 2026-09-13T00:00:00-07:00

window: [2026-08-30T07:00:00Z, 2026-09-13T07:00:00Z)
evidence: logical calls
grouped by: requested_policy
period     key                              calls   fail%  fallback%  retry%  timeout%  p50 sec  p95 sec         cost unpriced
---------- ------------------------------ ------- ------- ---------- ------- --------- -------- -------- ------------ --------
all        extract.bulk                     4,478    0.0%    1.2%    2.0%      0.1%   11.427   56.591 $  3.677362        0
all        extract.standard                 1,751    0.0%    1.6%    1.9%      0.1%    1.592   23.396 $  1.569982        0
all        image.default                    1,575    0.0%    2.2%    0.0%      0.0%    7.846   29.504 $ 31.670000        0
$ python -m creAItive.llm report fallbacks \
  --log-dir logs \
  --start 2026-08-30T00:00:00-07:00 \
  --end 2026-09-13T00:00:00-07:00

window: [2026-08-30T07:00:00Z, 2026-09-13T07:00:00Z)
evidence: fallback transitions
grouped by: reason
key                                        fallbacks    calls target response ok% calls recovered%
rate_limit                                      552      552              100.0%           100.0%
content_filter                                   49       49              100.0%           100.0%
timeout                                          20       20              100.0%           100.0%
provider_5xx                                     10       10              100.0%           100.0%

August 30–September 12, 2026. Overview excerpt: three policies shown. Fallback report covers all policies. Names anonymized; measurements unchanged.

03 / Learn from everyday use

Learn whether it helped.
Without another survey.

People already tell you something through how they use your app: they keep an answer, edit it, ask for another option, or undo a change. Those actions can provide useful quality evidence without interrupting their work.

A few annotations connect those actions to the model results behind them. creAItive.llm organizes this implicit feedback alongside explicit ratings and comments, keeping the execution history attached. You define what the signals mean in your product.

A product action becomes quality evidencePython
runtime.log_impact(
    answer_result,
    subject_ref=saved_answer_ref,
)

# When the user chooses to use this answer.
runtime.log_feedback(
    feedback_signal=answer_usage,
    action=llm.FeedbackAction.ACCEPTED,
    subject_ref=saved_answer_ref,
)
The application records a choice the user was already making. The shared reference connects it to the model call for quality analysis.
See sample output Two weeks of real use

Scroll right for more

$ python -m creAItive.llm quality-review --by requested_policy --limit 3 --limit-notes 0 \
  --log-dir logs \
  --start 2026-08-30T00:00:00-07:00 \
  --end 2026-09-13T00:00:00-07:00

recorded outcome actions by requested policy
signal                           version  group                                action         events   share
-------------------------------- -------- ------------------------------------ ------------ -------- -------
approach.choice                  1        assistant.advanced                   accepted          392   90.1%
approach.choice                  1        assistant.advanced                   rejected           21    4.8%
approach.choice                  1        assistant.advanced                   refined            19    4.4%

August 30–September 12, 2026. Excerpt: three recorded actions for one feedback signal. Names anonymized; counts and shares unchanged.

04 / Make improvement a repeatable process

From an observation
to a change you can assess.

Use operational and quality evidence to choose something worth improving. Capture representative cases and evaluate a new model or prompt before deploying it. After you ship, compare what happened in use.

The library supports that change lifecycle with evaluation tools, versioned execution records, and reports that compare evidence before and after a change. You and your agent decide what to ship through your application’s deployment process.

Evaluate first. Follow the change in use.CLI
# Summarize a completed, graded evaluation.
python -m creAItive.llm eval tabulate \
  evaluations/support-answer

# Inspect individual answers side by side.
python -m creAItive.llm eval compare \
  evaluations/support-answer

# Find recorded changes in the application.
python -m creAItive.llm changes list \
  --log-dir logs

# Compare the evidence around one change.
python -m creAItive.llm changes inspect \
  CHANGE_ID --log-dir logs --window 7d
Use your evaluation directory and a change ID from the report. Before-and-after evidence helps assess a change; it doesn’t establish cause on its own.

Illustrative excerpts from an application with its policies, credentials, and feedback definitions configured.

The companion book · Working draft

From a Model Call to an Intelligent System

The engineering practices behind reliable, improvable AI applications.

Geoff Hulten
Examples in Python using creAItive.llm

The thinking behind the code

The book explains why.
The library gives you a place to start.

Foundation models changed what one person can build. Working through coding agents, a builder can create software that interprets unfamiliar documents, exercises judgment, and takes actions in the world.

An LLM call can succeed mechanically and still fail in the way that matters: it didn’t help the person depending on it. This book follows the complete loop, from application intent through model execution to user outcomes and deliberate improvement.

It’s for programmers who haven’t operated a machine-learning system before. You’ll learn what to ask your coding agent to build, why the pieces fit together, and how to judge whether a change actually helped.

After years as an engineer, engineering leader, and teacher, I feel a little like a student again. This library grew out of building things that excited me. The book is my attempt to share the thinking behind it, so other people can build ambitious systems without having to rediscover all those lessons for themselves.

— Geoff Hulten, author of Building Intelligent Systems

Inside the book

  1. Welcome to LLM application development

    Explore what LLMs let your application do with language, documents, and other information. Decide which tasks need model judgment and which are better handled by ordinary code. Learn when to give an agent control over a sequence of work.

  2. Why creAItive.llm?

    Understand why model work benefits from a layer of its own. See how the library brings execution, reliability, and evidence together for you and your coding agent. Build the smallest complete implementation.

  3. Invoking model work

    Assemble requests, invoke the runtime, and use normalized results in your application. Connect each call to the application state it helped create. Establish the records you’ll need when you begin investigating operations and quality.

  4. Defining model execution

    Give the work stable names and decide which models and providers perform it. Configure settings, structured output, and recovery through policies and manifests. Keep those choices separate from the code that uses the result.

  5. Conversations and tool use

    Carry context across a conversation. Give a model controlled access to application capabilities through tools. Understand how the application executes tool requests and carries their results into the next turn.

  6. Monitoring your application

    Inspect cost, latency, errors, and recovery in a running application. Distinguish the call your application made from the provider attempts needed to complete it. Use reports to find changes in behavior and investigate individual results.

  7. Understanding quality

    A successful call can still produce an answer that fails the user. Learn how to assess answers when there isn’t one correct response and passing your tests doesn’t guarantee useful behavior. Distinguish whether a call completed, whether its answer was good, and whether it helped the user.

  8. Evaluating answers offline

    Capture representative requests and answers, then compare candidate models or prompts on the same work. Grade outputs without revealing which candidate produced them. Inspect the results and the important failures before deciding what to change.

  9. Learning from user outcomes

    Learn from what users keep, revise, replace, or undo without repeatedly asking them to rate the AI. Organize these implicit signals alongside explicit feedback and connect both to model results. Learn when a user’s action is useful evidence of quality and what context you need to interpret it.

  10. Connecting a custom model

    Decide when a local, hosted, or fine-tuned model is worth trying. Export production-derived examples for training and connect a compatible model through the existing execution layer. Use the same evaluation and operating practices to assess it.

  11. From a model call to an intelligent system

    Bring execution, operational evidence, user feedback, and evaluation into one improvement process. Follow a change from an observed problem through evaluation and back into use. Build an application you can maintain and improve over time.

Help shape the early version

Building with LLMs?
I’d like your take.

I’m looking for a few people to read the opening chapters or try the library on something they’re building. Especially programmers who haven’t spent years on an ML team.

Tell me where it helps, where I’ve skipped something, and where I’ve made it more complicated than it needs to be.

Email Geoff ↗

A sentence about what you’re building is plenty. Let me know if you’d like to read, try the code, or both.