Guides

Everything we have written down

Replacing an expensive model with a cheaper one is only worth doing if you can prove quality held. These are the pieces of that argument: how a prompt gets optimized, what makes a set of examples trustworthy, why prompts decay, and how to move off a platform that is closing. Every number in them comes from a run you can repeat.

Making a prompt better, provably

Start here if you have a prompt in production and no number that says whether it works.

Deciding what to adopt

Start here if you are choosing between tools, surfaces, or a way off a platform that is closing.

Or skip the reading and run the loop on one of your own tasks. Start a free pilot, or read the quickstart.

Apprentice
© 2026 Apprentice · Eval-gated LLM replacement