LLM evaluation: grading a model with no right answer
Vibes do not scale and BLEU measures nothing you care about. Eval sets, LLM-as-judge without fooling yourself, and the biases that corrupt it.
Calling models from Python, forcing structured output, choosing between fine-tuning and retrieval, and grading results.
Vibes do not scale and BLEU measures nothing you care about. Eval sets, LLM-as-judge without fooling yourself, and the biases that corrupt it.
Fine-tuning teaches behaviour, RAG supplies facts. A decision framework, the LoRA maths, real costs, and the cheaper ladder to climb first.
Asking a model to reply in JSON works until it does not. Learn why tool schemas beat prompt begging, and how to validate every response before it reaches your code.
BeginnerSend a prompt to an AI model from Python and print the reply — a short, beginner-friendly walkthrough with almost no setup drama.
Newsletter
Free essays on learning, Python, and ML — no account required. We’ll only email when there’s something worth reading.