Recommended models and effort settings for saving tokens #474
|
Which is the cheapest model and effort level that still works well for this course? I'd like to learn without burning tokens on a model that's bigger than the task needs. I see two use cases:
If you can, please share:
|
Replies: 3 comments 1 reply
|
I've been going through it with Claude Code, so here's what's worked for me:
Claude Sonnet, medium effort. That's plenty for explaining lessons, quizzes and recall warm-ups. Most of the work is reading the lesson and asking you questions, not deep reasoning.
Phases 0–10 barely need an API. It's mostly NumPy/PyTorch from scratch, so no tokens burned there. Rule of thumb: start cheap, and only move up when the failure is clearly the model and not your implementation. |
|
@kiril-buga @TomasD1az I'd be curious to see this measured as cost per successfully completed lesson, including all the rescue turns. A small model can look wonderfully cheap until one incorrect derivation sends you into five rounds of debugging the tutor. Congratulations, you've accidentally enrolled in an extra course. The repo's I'd compare a few model/effort combinations across equivalent lessons and record total token usage, retries, turns to completion, and code/test success. Then give the learner a fresh variation of the problem to solve unaided. Repeating the same quiz would contaminate the comparison as the learner becomes familiar with the answers. I'd also keep subscription-based tutors and API-backed lessons separate. Subscription usage limits don't translate cleanly into API dollars, whereas API labs can be compared using fixed inputs and cost per passing run, including failed attempts. @kiril-buga, would a small community results template be useful here? I'd be interested in seeing where the cheapest model actually stays cheap once retries enter the equation |
|
The first answer matches how the course is built, and measuring cost per completed lesson is the right next step. Some specifics from the repo: Lesson code. Only three lessons make a live provider call:
Every other lesson runs offline: NumPy or PyTorch from scratch, or a simulated client, so you learn the pattern without a key. CI runs the lesson tests with no API keys configured. All three read LLM_MODEL=claude-haiku-4-5 python3 phases/00-setup-and-tooling/04-apis-and-keys/code/first_api_call.py
LLM_MODEL=gpt-4o-mini python3 phases/11-llm-engineering/02-few-shot-cot/code/advanced_prompting.pyThe OpenAI client also reads OPENAI_BASE_URL=http://localhost:11434/v1 OPENAI_API_KEY=ollama LLM_MODEL=qwen3:8b python3 phases/11-llm-engineering/02-few-shot-cot/code/advanced_prompting.pyTutor. The skills ( Template. A results template would help. Keep the first version short: tool and model, effort, lesson path, total tokens or cost, turns to completion, retries, whether the code and quiz passed, and whether you then solved a fresh variation unaided. Post results as replies here so they stay comparable. |
I've been going through it with Claude Code, so here's what's worked for me:
Claude Sonnet, medium effort. That's plenty for explaining lessons, quizzes and recall warm-ups. Most of the work is reading the lesson and asking you questions, not deep reasoning.
For the math-heavy phases (1, 3, 7, 10: backprop, attention, building the LLM from scratch) I sometimes raise effort to high when I get stuck on a derivation. I switch back after.
Haiku works for quick lookups with course-guide, but as a tutor it felt too shallow. It tends to give you the answer instead of guiding you.
For certification prep I'd stay on Sonnet. The practice questions are bett…