In my previous article, Local Vibe Coding with Gemma 4 12B, I showed you how you can use an open source, open weights model such as Gemma 12B and do vibe coding locally. Your data stays within your boundaries, you don’t pay for tokens, and yet you can do useful work, locally. All you pay for is a powerful laptop and electricity.
While I showed vibe coding, you could use Gemma or any other similar models for other tasks such as image generation, document search, or anything you can throw AI at. And while I ran everything on a commercial-off-the-shelf MacBook Pro, my responses took a couple of seconds for most things, nothing stops you from setting up a server cluster and running a mini data center within your organization's walls, and effectively run useful AI, completely within your walls.
The value proposition is unquestionable. Your data stays with you. Frontier labs don't steal your data or business model. You don't pay for tokens as you go. In fact, I have found it quite questionable how we pay for AI, and I am sure I am not alone.
Let's say I have a budget of $5,000 to do a project. I am halfway through, and I burn through the budget because Opus 5 or GPT 5.6 decided my questions were complex enough to route to complex models, or I forgot to clear my sessions, or AI just made mistakes, and now I've burnt through $5,000. I'm stuck with a half-done project? Tell me one company CFO that would approve a project with an open-ended cost?
I am 100% sure the future belongs to local AI, followed by in-org AI on org-hosted clusters, followed by punting to the cloud when absolutely necessary.
Gosh I am being chatty, but there is so much more to talk about here, such as the Chinese models like Kimi K3 distilling Anthropic or Anthropic distilling OpenAI distilling New York Times or keep going back how Microsoft Word distilled WordStar to create compatible docs, and so on.
Such an exciting time to be a techie.
But let's refocus and level set what this article is all about.
What Is this Article About?
In this article, I am going to talk about distillation. It's funny, when I was younger, I talked about a completely different kind of distillation, which also got me in trouble. But this time it's different.
Model distillation (or knowledge distillation) is a technique where you train a smaller, faster model (the student) to mimic the behavior and outputs of a larger, more powerful model (the teacher).
Instead of training a model from scratch on raw, unrefined data, distillation transfers the “compressed knowledge” of a complex AI into a lightweight one that can run cheaply on everyday hardware or edge devices.
The advantage is that the student model is much more lightweight and usually specific to your needs. As a result, it can run on much lower end hardware at much lower cost. This opens the doors for local AI even wider. No doubt Opus or Mythos or Fable are incredibly smart, but they are overkill for fixing a Word document and they don't understand your organization's context. Extrapolate this to code, your company probably uses internal standards, libraries, shortcuts, that make no sense to frontier lab models. By having your distilled model, locally, you can run it at much lower cost, and your data stays yours.
So, distillation gives you drastic latency reduction, drastic RAM reduction, huge cost efficiency, and lets you get specialized performance. Sounds like a win-win situation.
Think of distillation like an experienced professor mentoring a student. The professor has read millions of books and understands complex nuances. Rather than making the student read every single book from scratch, the professor explains the concepts directly, provides curated examples, and scores the student's reasoning.
There are two main ways to do distillation.
Logit distillation, or probability distributions, refers to how a large teacher model makes a prediction. It doesn't just give one answer—it outputs a probability score across thousands of possible options (logits). An example of this would be, if shown a photo of a dog, a teacher model might say: Dog: 85%, Wolf: 10%, Car: 0.0001%. That 10% Wolf carries huge information: it tells the student model that dogs and wolves look similar, while cars do not. The student learns these subtle visual or logical relationships much faster than learning from simple pass/fail ground truth.
The other kind is data sequence distillation, or synthetic data, where the teacher model generates high-quality training examples—detailed reasoning steps, structured JSON, or clean code—that the smaller student model is trained on directly. The smaller student model is then trained directly on these teacher-generated responses.
I wanted to come up with a good example to demonstrate in this article, and I wanted to leave you with something realistic. Let me start with the onset here, everything I am talking about here, can work on “heavier” models like Gemma 12B. In fact, if you repeated these same steps with a heavier more capable model, you could end up with a specialized model that is truly impressive.
But those steps would take a few hours to run on my M4 Max MacBook Pro. To make things a bit more realistic, I decided to build a lightweight utility to parse terminal outputs. I ran into a classic problem: large models are too slow to run in the background, and small models aren't smart enough out of the box. This article and my example here is the exact playbook I use to solve that paradox.
So, let's set the stage. I am building a background terminal assistant, something that watches my Git logs and terminal outputs. When Git screams “fatal: refusing to merge unrelated histories”, I want my little assistant to quietly whisper, “Hey, you're trying to merge a totally different repository into this one. Use the --allow-unrelated-histories flag.”
Understood.
If I use a tiny 0.5B- or 1B-parameter model, it runs in milliseconds. It barely sips RAM. But if I ask it to explain a complex rebase conflict, it usually hallucinates or gives me a generic, useless response.
So, I am going to use knowledge distillation. I will show you how to take the deep, specialized knowledge of a massive “teacher” model and cram it into the tiny, lightning-fast brain of a “student” model. Best of all? I have designed this pipeline to be 100% operating-system agnostic. Whether you are on Windows, Linux, or a Mac, this will work seamlessly because we are splitting the workload: local data generation and cloud-based training.
What Are We Going to Build Here?
We are going to build, “The Git Whisperer”, a model that simplifies complex Git and terminal error messages into one-sentence explanations. In our example, the teacher will be Llama 3.2 3B (or Gemma 2 9B), running entirely locally using Ollama. The student will be Llama 3.2 1B Instruct, which is fast, tiny, but needs a good education.
Our classroom will be Google Colab (free tier). We will use Unsloth to train our model on a free cloud GPU so we don't have to wrestle with local hardware limits or CUDA drivers. You could technically follow this article on an iPhone if you're dedicated enough and have a keyboard, mouse, and a decent display.
Ready for the line to process.
Set Up the Teacher
First, we need our teacher model up and running. Ollama is the absolute easiest way to run local LLMs. It packages everything into a single executable. To do so, go to ollama.com/download (https://ollama.com/download) and grab the installer for your OS. When installed, go ahead and pull down the teacher model by running the following in terminal.
ollama run llama3.2
This command downloads the Llama 3.2 3B model (around 2GB) and starts an interactive chat. You can type a message to verify it works, then type /bye to exit. Ollama is now running as a background service on your machine, ready to receive API calls.
Generate the Synthetic Dataset
We need to create a textbook for our student. To do this, I will write a Python script that hands hundreds of terrible, confusing Git errors to the teacher model and asks it to explain them simply. I will use Python to create my synthetic dataset. To do so, go ahead and create a directory on your disk, open it in VSCode, and set up a .venv (virtual environment), and install the Python library using requirements.txt. I realize I am fast forwarding a few Python basics here, but a simple Google search will sort it out for you. But to simplify, essentially you are running this:
pip install ollama
Now, create a file called generate_data.py. It generates our synthetic data. Let's build it, bit by bit.
Starting with Listing 1. Do note that in Listing 1, to make things readable, I have truncated some of the text strings. The code sets up the foundation for generating synthetic training data using a local language model. You are preparing a list of confusing Git error messages and defining instructions for an AI to translate them into plain, beginner-friendly terms.
Listing 1: The setup and prompts
import json
import ollama
# We define exactly how we want the model to behave
system_prompt = "You are a helpful mentor. ..."
# A sample of raw, ugly errors.
# (In reality, you'd want 100-500 of these)
raw_errors = [
"fatal: Rebase workflow failed due ..",
"error: failed to push some refs to 'origin'..",
]
The system_prompt is intended to be, “You are a helpful mentor. Translate complex Git and technical error messages into simple, one-sentence explanations for beginners.” This is how we are strictly constraining the teacher. We don't want paragraphs; we want exactly one sentence. This constraint is what the student will eventually learn.
Then we have a bunch of raw_error messages. These are the confusing, hard-to-understand terminal errors that Git throws at us. Such errors make me scratch my head sometimes, and that I want my student to translate into plain English for me.
Next let's focus on the generation loop, shown in Listing 2. The code creates an empty list called dataset to hold the finished training examples and prints a quick update to the screen, so you know the generation process has started. It then loops through every Git error in the list you provided earlier, sending each error one at a time to your local Llama 3.2 model via the Ollama tool. For each request, it passes along both your system instructions, telling the model to act as a mentor, and the specific error message as the user's input. Once Llama 3.2 generates a response, the script extracts the text, cleans up any leading or trailing whitespace, and prints out a short progress message showing the first 30 characters of the error being processed. Finally, it packages the system instruction, the original error message, and the model's simplified explanation into a structured chat format and saves it as an entry in the dataset list, preparing the data for model training.
Listing 2: The generation loop
dataset = []
print("Generating teacher responses...")
for error in raw_errors:
# Call the local Ollama API
response = ollama.chat(
model='llama3.2',
messages=[
{'role': 'system', 'content': system_prompt},
{'role': 'user', 'content': error}
]
)
teacher_answer = response['message']['content'].strip()
print(f"Processed: {error[:30]}...")
# We save this in the standard JSONL format
# expected by training libraries
dataset.append({
"messages": [
{"role": "system", "content": system_prompt},
{"role": "user", "content": error},
{"role": "assistant", "content": teacher_answer}
]
})
Nicely done! The hard work is done; now we focus on saving the output, as shown in Listing 3.
Listing 3: Saving the data
# Write the list of dictionaries to a JSON Lines file
with open("train.jsonl", "w") as f:
for entry in dataset:
f.write(json.dumps(entry) + " ")
print("Dataset generated successfully as 'train.jsonl'!")
Now go ahead and run this script. Ensure you have Ollama running and Llama 3.2 downloaded. This script will chug away for a minute using your local machine's resources and pop out a beautiful train.jsonl file. This file is the gold we need. It contains the perfect inputs and the perfect outputs.
One thing I'll add. You could use AI to greatly enhance the raw_errors. Really all you need is a dump of all possible confusing Git messages.
Establishing a Baseline
To keep things consumable in a simple article, I have deliberately picked a lightweight teacher model and very few messages. It makes sense to get an idea of how bad our student is before we teach it. If you were to run the raw Llama 3.2 1B model and give it our system prompt and a Git error, you would typically see one of three failures.
You could see the rambler, where it ignores the “1-sentence” rule and writes a three-paragraph essay on the history of Linus Torvalds creating Git.
Or you could see the hallucinator, which suggests running a command that doesn't exist, like “git fix rebase –auto-magic” or some nonsense like that.
Or you might see the parrot, which just repeats the error message back to you.
After three decades in this field of work, I still find all these funny. Anyway, we need to fix all these. But I encourage you to try it. Fire up Ollama, and give it a raw command, and have it explain what it thinks that command is.
I tried it, with the following prompt
*What does this error mean? “error: failed to push some refs to ‘origin’ updates were rejected because the tip of your current branch is behind”
*Figure 1 shows the results. I had to cut it a bit short because it literally gave me pages of text.

The Classroom (Google Colab)
Now, I will show you how to train the student. We are leaving our local machine and going to the cloud. Why? Because fine-tuning requires a GPU with decent VRAM, and dealing with CUDA errors on Windows or PyTorch architecture bugs is a nightmare for beginners. You could do it, but then my article would also be Mac- or Windows-specific. So, we'll use the free tier of Google Colab.
Start by setting up the notebook. Go to colab.research.google.com and create a New Notebook. Go to Runtime > Change runtime type and select T4 GPU. This can be seen in Figure 2.

Next, on the left-hand side, look for an icon that looks like a folder. Its tooltip should say “Files”. Click it and drag and drop the training.jsonl file you had created earlier.
With our data there, let's start writing our notebook. The first step is to install dependencies. This can be seen in Listing 4.
Listing 4: Install dependencies
!pip install "unsloth[colab-new]
@ git+https://github.com/unslothai/unsloth.git"
!pip install --no-deps x
formers trl peft accelerate bitsandbytes
Next, we load our student model shown in Listing 5. This code sets up a lightweight language model for efficient fine-tuning using the Unsloth library and PyTorch. It starts by defining configuration settings, including a maximum text window of 2,048 tokens, automatic data type detection, and 4-bit quantization to drastically reduce memory usage. Next, it downloads and loads the pre-trained Llama 3.2 1-billion Instruct model along with its tokenizer. Finally, instead of training the entire multi-billion parameter model, it attaches low-rank adaptation (LoRA) adapters directly to key attention layers (q_proj, k_proj, v_proj, and o_proj). This attaches a small, trainable “adapter layer” to the model so you can teach it new behaviors in minutes without needing massive hardware.
Listing 5: Load the student model
from unsloth import FastLanguageModel
import torch
max_seq_length = 2048
dtype = None # Auto-detect
load_in_4bit = True # Memory saving technique
# Load the 1B parameter student model
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Llama-3.2-1B-Instruct",
max_seq_length=max_seq_length,
dtype=dtype,
load_in_4bit=load_in_4bit,
)
# Add LoRA adapters (this is what we are actually training)
model = FastLanguageModel.get_peft_model(
model,
r=16,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
lora_alpha=16,
lora_dropout=0,
bias="none",
)
Instead of retraining all one billion parameters of the model which would take weeks, we are using a technique called LoRA (low-rank adaptation). We are essentially taping a small, temporary brain to the side of the model, training only that small tape, and letting it steer the big brain. This is why we can do this in minutes on a free GPU.
Next, let's format the dataset. The code shown in Listing 6 prepares your synthetic dataset so the Llama 3 model can learn from it. First, it updates your model's tokenizer using Unsloth's chat template tool, ensuring the data will be wrapped in the exact special control tokens Llama 3 expects to distinguish between system instructions, user inputs, and model responses. It then defines a helper function, formatting_prompts_func, which takes raw message histories and converts each structured conversation into a single, properly formatted string without turning them into numerical tokens just yet. Finally, the script loads your uploaded train.jsonl file into a Hugging Face dataset object and runs every example through that helper function in batches, outputting a clean, formatted dataset ready to feed directly into the trainer.
Listing 6: Format the dataset
from datasets import load_dataset
# Unsloth provides a handy chat template mapper
from unsloth.chat_templates import get_chat_template
tokenizer = get_chat_template(
tokenizer,
chat_template="llama-3",
mapping={
"role": "role",
"content": "content",
"user": "user",
"assistant": "assistant",
},
)
def formatting_prompts_func(examples):
texts = []
for messages in examples["messages"]:
# Apply the specific Llama 3 formatting tokens
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=False,
)
texts.append(text)
return {"text": texts}
# Load the JSONL file we generated locally
# and uploaded to Colab
dataset = load_dataset(
"json",
data_files="train.jsonl",
split="train",
)
dataset = dataset.map(
formatting_prompts_func,
batched=True,
)
Now finally, we can write the code to do the training, shown in Listing 7.
Listing 7: Training code
from trl import SFTTrainer
from transformers import TrainingArguments
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
train_dataset=dataset,
dataset_text_field="text",
max_seq_length=max_seq_length,
args=TrainingArguments(
per_device_train_batch_size=2,
gradient_accumulation_steps=4,
warmup_steps=5,
max_steps=60,
learning_rate=2e-4,
fp16=not torch.cuda.is_bf16_supported(),
bf16=torch.cuda.is_bf16_supported(),
logging_steps=1,
optim="adamw_8bit",
weight_decay=0.01,
lr_scheduler_type="linear",
seed=3407,
output_dir="outputs",
),
)
trainer_stats = trainer.train()
This code sets up and runs the actual training loop for your student model using Hugging Face's SFTTrainer, supervised fine-tuning trainer. It configures all the hyperparameters like batch sizes, learning rate, and memory optimizations and then kicks off the fine-tuning process.
First I import the necessary training tools. SFTTrainer simplifies fine-tuning language models on formatted text datasets, and TrainingArguments holds all the detailed settings that control how the model learns. The SFTTrainer(...) creates the trainer instance by passing in your model, tokenizer, train_dataset, and specifying text as the dataset column containing the formatted prompts.
We also pass in various TrainingArguments(...) settings as follows.
The per_device_train_batch_size = 2 and gradient_accumulation_steps = 4, instead of processing a huge batch of data all at once (which crashes VRAM), the trainer processes two examples at a time and accumulates gradients over four steps. This effectively simulates a larger batch size of eight ($2 \times 4$) while keeping GPU memory usage low.
The warmup_steps = 5, gently increases the learning rate during the first five training steps to prevent sudden drastic updates that could ruin the pre-trained weights.
The weights.max_steps = 60, tells the trainer to stop after 60 training steps which is perfect for a small dataset demo.
The learning_rate = 2e-4, sets the speed at which the model updates its weights.
fp16 / bf16 automatically checks your GPU's capabilities and enables 16-bit precision math to speed up training and save VRAM.
optim = “adamw_8bit”, uses an 8-bit quantized version of the AdamW optimizer to further reduce GPU memory footprint.
weight_decay = 0.01, is a regularization technique that prevents the model from overfitting to your training
data.seed = 3407, sets a fixed random seed so that your training results are reproducible every time you run it.
output_dir = “outputs”, simply saves checkpoints and logs to a directory named outputs.
And finally, trainer_stats = trainer.train() is the line that actually runs the training! It starts the loop, updates the LoRA weights step by step, prints loss metrics to your console/notebook, and returns a summary of the training performance once complete.
Now you might be wondering what it might take to run all this locally? The answer is, not much other than the code you already have. Except you'll be using your local GPU instead of Google's. But for now, click Run. You will see a beautiful progress bar appear. Watch the training loss go down. It might start at 2.5 and drop down to 0.8. This number is literally the model getting less confused by your dataset over time. Because the student model is only 1B parameters and Unsloth is highly optimized, this training process on a small dataset will finish in less than five minutes.
Testing the Model
While still in your notebook, verify that you see “checkpoint-60” created, as can be seen in Figure 3.

If you see that, training is done. The LoRA adapters are locked in. Let's see if the student actually learned to behave like the teacher.
So now, let's try it out. Let's test our model, also known as “inference”. Go ahead and place another cell in your notebook, and put the code as shown in Listing 8 in the inference cell. Note that I have truncated some strings for readability. The full system prompt is, “You are a helpful mentor. Translate complex Git and technical error messages into simple, one-sentence explanations for beginners.”
Listing 8: Inference cell
# Enable native inference speedups
FastLanguageModel.for_inference(model)
# 1. Define the system prompt used during training
system_prompt = "You are a helpful mentor. ..."
# 2. Set up a test prompt that the model has NEVER seen before
test_messages = [
{"role": "system", "content": system_prompt},
{
"role": "user",
"content": "fatal: Not a git repository (or any of the parent directories): .git"
}
]
input_text = tokenizer.apply_chat_template(
test_messages,
tokenize=False,
add_generation_prompt=True
)
inputs = tokenizer([input_text], return_tensors="pt").to("cuda")
# 3. Generate the response
outputs = model.generate(
**inputs,
max_new_tokens=64,
use_cache=True
)
response = tokenizer.batch_decode(outputs, skip_special_tokens=True)[0]
# Print only the assistant's response
print(response.split("assistant\n")[-1])
The code shown in Listing 8 tests your freshly fine-tuned student model on a brand-new Git error message it has never seen before. It formats the request, feeds it to the model on the GPU, and prints out only the model's simplified response to verify that the distillation was successful.
Using FastLanguageModel.for_inference(model), this prepares the model for testing (inference) mode rather than training mode. Unsloth applies specific performance optimizations here that make generation up to twice as fast.
The system_prompt sets the behavior instructions for the model, ensuring it remembers its persona as a mentor that explains technical errors in a single, simple sentence.
The test_messages packages the test scenario into a standard conversation structure containing both the system prompt and a real-world user query.
Then there is formatting and tokenization. The apply_chat_template method converts the structured test_messages list into a single formatted string wrapped in Llama 3's special control tokens, adding an assistant header at the end to prompt the model to answer. The tokenizer(...).to("cuda") statement converts that formatted text string into numbers (tokens) that the model understands and sends them directly to the GPU (cuda) for execution.
Then we run the model to predict the next words, allowing it to output up to 64 new tokens. use_cache=True speeds up the output generation by reusing previous calculations. And using batch_decode, we convert the model's numerical output tokens back into readable English text, stripping out special control symbols.
Finally, we print the response, as can be seen in Figure 4. I'll clear out the cruft for you, the message output is “This error message is telling you that your system is trying to interact with the git repository, but it's not configured to do so.”
Ready.

Export for Real World Usage
As impressive as this is, this is still running in a Google notebook. Your brilliant student, stuck in Google Colab, isn't going to do the world any good, is it? We want to bring it back home to your local machine so you can use it in your actual terminal utility. Unsloth makes it incredibly easy to merge the LoRA adapters into the base model and export it as a .gguf file , which surprise surprise, is the exact format Ollama uses!
To achieve this, add another cell in your notebook and add the code shown below in that cell, and execute it. Once this is finished, look in your Colab files panel. You will see a file called llama-3.2-1b-instruct.Q4_K_M.gguf. Download this file to your computer. It will be tiny—around 800MB.
model.save_pretrained_gguf(
"model",
tokenizer,
quantization_method="q4_k_m"
)
This can be seen in Figure 5.

This is huge, I mean not literally but figuratively. We have turned a multi-gigabyte model into a specialized tiny model at around 700 MB using minutes of distillation. Not bad.
Using the Model Locally
Once the model is exported to the Colab files panel, go ahead and download it to your local computer. Place this file in a new folder. In that same folder, create a file named Modelfile (no extension) in the same directory as your downloaded .gguf file. Add the text seen in snippet below to the Modelfile file.
FROM ./llama-3.2-1b-instruct.Q4_K_M.gguf
SYSTEM "You are a helpful mentor. ..."
Please note again that I have truncated the system prompt. It's that system prompt we have been using all through this article.
Now fire up terminal and go to that directory where you have placed the Modelfile and the gguf file. Once in that folder, run the following commands:
ollama create git-whisperer -f Modelfile
ollama run git-whisperer
The output of creating git-whisperer can be seen in Figure 6.

Once you run “ollama run git-whisperer”, you should see a prompt, where you can start interacting with this new distilled model. You did it. You successfully distilled a large model's specialized knowledge into a lightweight, sub-gigabyte model that responds instantly on your local hardware.
Next, I went online, and Googled for “complex git error message” and pasted that in. Figure 7 shows what my output looked like.

To be specific, it took my complex Git error message and gave me a simple output for it as follows.
This error message is telling you that your local version of some files on your GitHub account isn't up to date with the original versions, and if you try to push those files to GitHub, they will lose their changes.
jaw.pickFromFloor()++; // :-O
Summary
I often find myself trying to learn a complex topic and getting overwhelmed by complicated language. Learning any new thing is like trying to bite into a big burger. You know it's gonna be tasty, but you also know you are going to make a mess of it. So far my approach has been to just dive in and figure it out as I go. But do I wish I had a mentor that I can tap on the shoulder at any point and ask for help, “Hey buddy, what does this mean, break it down for me at dummy level, please?”
Imagine, you are a computer engineer, trying to learn about investing. And the prose you are reading talks about WACC, weighted average cost of capital. What the heck does that even mean? Don't worry, just distill a model and you are golden.
Now I don't think you'll do that. Most likely you'll just Google for that term and party on. But expand your imagination a bit. You are a company and you have lots of complex internal tooling. You want a helpful buddy to answer questions that are specific to your company. And you want to do so using local hardware and not send your proprietary data outside your walls.
This is where local distillation comes in. In this simple article, you broke this process down by using Ollama locally for data generation and Google Colab for training, completely removing the operating system and hardware barriers. You used AI to create synthetic data; you then uploaded that to a Google Colab training space where you used no local compute and for free distilled a fairly capable model that can explain complex Git commands easily and consuming almost no local resources.
Here is a little project for you if you are up for it. Using exactly the same concepts here, can you distill a model that runs Git commands for you? Try using Gemma 12B as a starting point, and on my powerful M4 Max, it took me a couple of hours to distill this. Create a gh helper, and now you can interact with Git, using plain English commands.
Try it!
Please provide the line to process.



