In part 1 on fine-tuning large language models (LLMs) (https://www.codemag.com/Article/2610031/Fine-Tuning-Large-Language-Models-Part-1), I explored the fundamental approaches to adapting a pretrained model, from feature extraction and partial fine-tuning to full fine-tuning and LoRA. Part 2 takes the next step, going deeper into fine-tuning techniques while also looking at what happens after training.

We begin with quantization, a technique for reducing model size and memory requirements that makes LoRA fine-tuning practical even on hardware that cannot accommodate a full-precision model. We then explore prompt tuning, another lightweight alternative that adapts a model without modifying most of its parameters.

But fine-tuning a model is only half the story. How do you know whether it actually improved? This part examines how to evaluate your results using standard metrics, baseline comparisons, and checks for catastrophic forgetting. For tasks where there is no single correct answer, we'll also look at LLM-as-a-judge evaluation.

Finally, we move from training to deployment. Knowledge distillation shows how a large, capable model can be used to create a smaller and more efficient model. We then look at exporting and deploying fine-tuned models using ONNX, regardless of the platform used for training.

By the end of this article, you'll have a broader view of the fine-tuning lifecycle—not just how to adapt a model, but how to compress it, evaluate it, and ultimately ship it.

Quantizing a Model

In the previous article, I discussed several fine-tuning techniques, including full fine-tuning and parameter-efficient fine-tuning (PEFT) methods such as low-rank adaptation (LoRA). When working with larger models, memory becomes the binding constraint even for LoRA, because the frozen base model's weights still need to be loaded into memory in full precision. Quantization addresses this by loading the frozen base model at reduced numerical precision (commonly 4-bit) while keeping the small trainable LoRA adapters at full precision—this combination is what's generally referred to as QLoRA.

A Beginner's Explanation of Quantization

Model weights are normally stored as 32-bit or 16-bit floating point numbers. Quantization compresses these down to lower-precision representations—8-bit or even 4-bit—trading a small amount of numerical precision for a large reduction in memory footprint (roughly 4x smaller going from 16-bit to 4-bit). Because the frozen base model in a LoRA setup is only ever read from, not updated, it tolerates this precision reduction far better than a model you're directly training would.

At the time of writing, the most widely used 4-bit/8-bit quantization library, bitsandbytes, has its most mature, fully supported path on NVIDIA GPUs with CUDA. The library has been going through a multi-backend rewrite, though, and recent releases also ship experimental support for Apple Silicon and AMD ROCm—so the code in this article may now run directly on a Mac, just on a newer, less battle-tested code path than the CUDA one. If it doesn't work for you, or you'd rather use a more established Apple Silicon option, two fallbacks remain: use LoRA without 4-bit quantization, which ordinary system RAM or a laptop's unified memory already handles comfortably for small-to-mid-sized models, or use Apple's own MLX framework, which has native quantization support built specifically for Apple Silicon. Either way, if you have an NVIDIA GPU, the bitsandbytes path remains the most reliable choice.

Which code to use here depends on your hardware—the following sections show the two options available.

On an NVIDIA GPU (CUDA)—or Experimentally on Apple Silicon—via bitsandbytes

Listing 1 shows how to set up QLoRA fine-tuning for the Qwen2.5-7B-Instruct model—loading it in 4-bit quantized precision and preparing it for parameter-efficient training with LoRA adapters.

Listing 1: Loading Qwen2.5-7B-Instruct with 4-bit quantization and configuring LoRA for fine tuning

from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
    BitsAndBytesConfig,
)
from peft import (
    LoraConfig,
    get_peft_model,
    prepare_model_for_kbit_training,
    TaskType,
)
import torch
model_name = "Qwen/Qwen2.5-7B-Instruct"
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    quantization_config=bnb_config,
    device_map="auto",
)
model = prepare_model_for_kbit_training(model)
lora_config = LoraConfig(
    task_type=TaskType.CAUSAL_LM,
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    target_modules=[
        "q_proj", "v_proj", "k_proj", "o_proj",
    ],
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()

The prepare_model_for_kbit_training() function handles several important details automatically: it casts certain layers to a stable precision for training, enables gradient checkpointing to further reduce memory, and ensures the model is in the right state for gradients to flow correctly into the LoRA adapters despite the frozen base being quantized.

Understood.

  • 4-bit quantization—BitsAndBytesConfig loads the model in NF4 4-bit format, computes in bfloat16 for stability, and applies double quantization to shrink memory further. This lets a 7B model fit on a single consumer GPU.
  • Model/tokenizer loading—Loads both from Hugging Face Hub; device_map="auto" auto-distributes layers across available devices.
  • k-bit training prep—prepare_model_for_kbit_training() adjusts the quantized model for training (e.g., casting norms to float32, enabling gradient checkpointing).
  • LoRA config—Defines low-rank adapters (r=16, alpha=32, dropout=0.05) applied only to attention projections (q/k/v/o_proj), keeping the rest of the model frozen.
  • Apply LoRA—get_peft_model() injects the adapters; print_trainable_parameters() shows the trainable fraction is typically under 1% of total parameters.

The net effect: combines 4-bit quantized weights with LoRA adapters to fine-tune a 7B model cheaply, updating only a small set of parameters while the base model stays frozen.

On Apple Silicon, via mlx-lm

The Apple Silicon path through bitsandbytes is new and still labeled experimental—expect rougher edges, missing operations, or a silent fallback to a slower, less-optimized kernel than the CUDA path gets. If it doesn't work cleanly on your setup, or you want a more mature Apple Silicon option, MLX below is the better-established choice.

To use mlx-lm on your Mac, install the following packages:

!pip install "transformers>=5.7.0,<5.13.0"
!pip install mlx-lm

The following code snippet loads a community-published 4-bit MLX checkpoint directly—MLX quantizes ahead of time, at conversion, rather than at load time the way bitsandbytes does.

from mlx_lm import load

model, tokenizer = load(
    "mlx-community/Qwen2.5-7B-Instruct-4bit"
)

Understood.

  • from mlx_lm import load** imports the loading utility from mlx-lm, a library built on top of MLX specifically for running LLMs efficiently on Mac hardware (leveraging the unified memory architecture of M1/M2/M3/M4 chips).
  • load(“mlx-community/Qwen2.5-7B-Instruct-4bit”) downloads (if not cached) and loads a pre-quantized 4-bit version of Qwen2.5-7B-Instruct from the mlx-community organization on Hugging Face Hub—a community that maintains MLX-converted versions of popular models.
  • The function returns two objects: model: the loaded MLX model, ready for inference (or further fine-tuning via mlx-lm's LoRA utilities). tokenizer: the corresponding tokenizer for encoding/decoding text.

Prompt-Based Fine-Tuning

As language models have grown into the tens or hundreds of billions of parameters, full fine-tuning—updating every weight in the model—has become increasingly expensive and often impractical. Prompt-based fine-tuning offers a lightweight alternative: instead of touching the model's weights at all, it adapts the model's behavior by manipulating what gets fed into it. The frozen model stays exactly as it was pretrained, while a small, task-specific prompt does the work of steering it toward the desired output.

This family of techniques spans a spectrum from fully human-readable instructions to entirely learned, non-interpretable vectors. Understanding where a given method sits on that spectrum—and what's actually being optimized—is key to understanding how and why these approaches work.

Soft Prompts Vs. Hard Prompts

Ordinary prompting—typing an instruction like “Translate this to French:” ahead of your input—uses what's sometimes called a hard prompt: human-readable text that you write and the model interprets.

Prompt tuning instead learns a soft prompt: a short sequence of continuous embedding vectors that have no corresponding words at all. They aren't human-readable text you could type—they're numbers, discovered by gradient descent, that get prepended to your actual input's embeddings before the (entirely frozen) model processes them.

Think of a soft prompt as a learned “steering wheel”: it doesn't change a single weight inside the model, but it reshapes the effective input in a way that reliably nudges the frozen model's behavior toward your task.

Figure 1 shows the conceptual idea of using soft prompts and hard prompts to send a question to a model. Note that soft prompts don't necessarily have to be paired with a hard prompt. Pure prompt tuning can work with just the virtual tokens preceding the raw input (no instructional text at all).

Figure 1: Soft prompts (invisible, learned) and hard prompts (visible text) are concatenated and fed into a frozen model
Figure 1: Soft prompts (invisible, learned) and hard prompts (visible text) are concatenated and fed into a frozen model

What Gets Trained

This is the most parameter-efficient technique in this article. The base model is entirely frozen—every weight, every attention layer, every embedding stays exactly as pretrained. The only thing trained is a small set of virtual token embeddings, often just a few dozen to a few hundred vectors, which is why prompt tuning can require orders of magnitude fewer trainable parameters than even LoRA.

Listing 2 shows how to set up prompt tuning using Hugging Face's peft library. Instead of injecting LoRA adapters into the model's layers, this approach freezes the entire base model and learns only a small set of virtual token embeddings that get prepended to every input.

Listing 2: Setting up a prompt-tuning model with Hugging Face peft

from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
)
from peft import (
    PromptTuningConfig,
    PromptTuningInit,
    get_peft_model,
    TaskType,
)
model_name = "Qwen/Qwen2.5-1.5B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_name)
base_model = \
  AutoModelForCausalLM.from_pretrained(model_name)
prompt_config = PromptTuningConfig(
    task_type=TaskType.CAUSAL_LM,
    prompt_tuning_init=PromptTuningInit.RANDOM,
    num_virtual_tokens=20,  # embedding vectors to 
                            # prepend
    tokenizer_name_or_path=model_name,
)
model = get_peft_model(base_model, prompt_config)
model.print_trainable_parameters()

The key differences from the earlier LoRA setup are:

  • No quantization needed—since prompt tuning trains far fewer parameters than LoRA, a 1.5B model can be loaded in full precision without BitsAndBytesConfig.
  • PromptTuningConfig replaces LoraConfig. prompt_tuning_init=PromptTuningInit.RANDOM initializes the virtual token embeddings with random values (an alternative is PromptTuningInit.TEXT, which seeds them from the embeddings of an actual phrase you provide—often giving faster convergence).
  • num_virtual_tokens=20 sets how many learned embedding vectors will be prepended to the input sequence. This is the entire trainable “soft prompt”—typically anywhere from a handful to a few hundred tokens, depending on task complexity.
  • tokenizer_name_or_path is required so the library can determine the model's embedding dimension, ensuring the virtual token vectors match the size the model expects.

As before, get_peft_model() wraps the base model with this configuration, and print_trainable_parameters() reveals just how small the trainable footprint is—with prompt tuning, it's often a tiny fraction even of what LoRA trains, since only the virtual token embeddings themselves are updated and every other parameter, including the model's own embedding table, stays frozen.

A typical result here trains on the order of tens of thousands of parameters—a tiny fraction even of LoRA's already-small footprint—because you're only ever learning num_virtual_tokens × embedding_dimension values, nothing more.

A Concrete Example: Steering Style

Suppose you fine-tune with soft prompts on a dataset of Yelp restaurant reviews. Before tuning, prompting the model with “I was thinking” might continue with something generic: “I was thinking of going to the store tomorrow…” After prompt tuning on review data, the same input, now with the learned virtual tokens silently prepended, might continue: “I was thinking this restaurant had great potential but the service was disappointing…” The model's weights never changed—only the invisible learned “context setter” in front of the input did, and that alone was enough to reliably shift the model's output into a Yelp-review style and domain.

Here's the same idea end-to-end, in code, continuing directly from the tokenizer, base_model, and prompt_config set up previously, using the full Yelp Review Full dataset (Hugging Face's Yelp/yelp_review_full) to actually steer the model the way described above.

Listing 3 shows the full training loop that fine-tunes the virtual token embeddings on the Yelp Review Full dataset, teaching the model to write in the style of a Yelp review.

Listing 3: Prompt-based fine-tuning with the full Yelp reviews dataset

import torch
from datasets import load_dataset
from transformers import (
    DataCollatorForLanguageModeling,
    Trainer,
    TrainingArguments,
)


def get_device():
    if torch.cuda.is_available():
        # Windows/Linux with NVIDIA GPU
        return torch.device("cuda")

    if torch.backends.mps.is_available():
        # Apple Silicon Mac
        return torch.device("mps")

    # Fallback, works everywhere
    return torch.device("cpu")


device = get_device()
base_model = base_model.to(device)

# Full Yelp Review Full dataset -- ~650,000 reviews
raw_dataset = load_dataset("Yelp/yelp_review_full")
train_dataset = raw_dataset["train"]

# subsample with a smaller dataset
train_dataset = train_dataset.shuffle(seed=42).select(range(2000))


def tokenize_function(examples):
    return tokenizer(examples["text"], truncation=True, max_length=64)


tokenized_train = train_dataset.map(
    tokenize_function,
    batched=True,
    remove_columns=["text", "label"],
)

data_collator = DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False)

# prompt_config
model = get_peft_model(base_model, prompt_config)

training_args = TrainingArguments(
    output_dir="./prompt-tuning-output",
    per_device_train_batch_size=8,
    num_train_epochs=1,  # 650K reviews is enough
    # for 1 pass
    learning_rate=3e-2,  # prompt tuning tolerates
    # a higher LR
    logging_steps=500,
    report_to="none",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_train,
    data_collator=data_collator,
)

trainer.train()

Understood.

  • get_device() picks the best available backend—CUDA on NVIDIA GPUs, MPS on Apple Silicon, or CPU as a fallback—making the script portable across machines without hardcoding a device.
  • Dataset and tokenization—the full ~650,000-review Yelp dataset is loaded and tokenized with max_length=64, keeping sequences short since the goal is to learn a style, not memorize long-form content. remove_columns=["text", "label"] strips the original fields once tokenized, since the label (star rating) isn't used here—only the review text matters for shaping the model's writing style.
  • DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False) configures the collator for causal (next-token prediction) language modeling rather than masked language modeling, and handles dynamic padding within each batch.
  • get_peft_model(base_model, prompt_config) applies the PromptTuningConfig from the previous listing, freezing the base model and exposing only the 20 virtual token embeddings as trainable parameters.
  • Training arguments—a single epoch is sufficient given the dataset's size, and the learning rate (3e-2) is notably higher than what's typical for full fine-tuning or even LoRA. This is characteristic of prompt tuning: since only a tiny number of parameters are being optimized, larger updates can be tolerated without destabilizing training.
  • trainer.train() kicks off the training run, updating only the virtual token embeddings while every other weight in the model—including the original embedding table—remains frozen throughout.

Training on the full 650,000-review split takes meaningfully longer than a handful of examples would—on a single epoch, expect anywhere from tens of minutes on a GPU to a few hours on CPU or Apple Silicon. If you just want to see the steering effect quickly, subsample first: train_dataset = train_dataset.shuffle(seed=42).select(range(2000)).

With training complete, compare the frozen base model's output to the prompt-tuned model's output on the exact same prompt, as shown in Listing 4.

Listing 4: Comparing base and fine-tuned model outputs with a generate function

def generate(
    active_model, prompt, max_new_tokens=150
):
    inputs = tokenizer(
        prompt, return_tensors="pt"
    ).to(device)
    active_model.eval()
    with torch.no_grad():
        output = active_model.generate(
            **inputs,
            max_new_tokens=max_new_tokens,
            do_sample=False,
        )
    return tokenizer.decode(
        output[0], skip_special_tokens=True
    )
print(
    "Before:",
    generate(base_model, "I was thinking"),
)
print(
    "After: ",
    generate(model, "I was thinking"),
)

When you run the above code, you can compare the output of the model before the fine-tuning:

I was thinking about the following
problem:
> Let $f$ be a function defined on
> $\mathbb{R}$ such that for all real
> numbers $x$, we have
>  $$ f(x) = \frac{x^2}{1+x^4} +
> \frac{1}{\sqrt{1-x}}. $$
Is it possible to find an explicit
formula for $f^{-1}(y)$, where $y$ is
any real number? If so, what would
this formula look like?
To determine if there exists an
explicit formula for \( f^{-1}(y) \),
we need to analyze the given function
\( f(x) = \frac{x^2}{1+x^4}

And this is the output of the model after the fine-tuning:

I was thinking about going to this
place for a while. I've heard good
things, but not enough that I would go
out of my way just because someone told
me it's great.
So when I saw the menu and thought "Oh,
I can get some decent food here!" I
decided to give it a try. The first
thing I noticed is how much better the
food looks than what you see in
pictures. It's like they have their own
secret recipe or something.
The second thing I noticed is how clean
everything is. They don't seem to be
using any disposable utensils or
anything like that. Everything is
served on nice plates with no messes
at all.
The third thing I noticed is how
friendly

Because this now trains on real, full-scale Yelp review text rather than a handful of hand-picked lines, the After output should pick up the Yelp-review voice and vocabulary far more reliably. Either way, base_model's own weights, and the vast majority of the prompt-tuned model's, never changed at all—only the twenty virtual token embeddings did.

When to Reach for Prompt Tuning

Prompt tuning is worth choosing when you need maximum parameter efficiency above all else—for rapid prototyping, for deploying many different task-specific behaviors cheaply on top of a single shared frozen model, or when your compute budget is extremely limited. Its ceiling on task performance is generally lower than LoRA's, and considerably lower than full fine-tuning's, since it can only ever influence the model indirectly through the input, never by adapting the model's internal computation itself. Treat it as the right tool when “good enough, extremely cheap, and easy to swap” beats “as strong as possible.”

Choosing Among the Three PEFT-Style Options

With prompt tuning, LoRA, and full fine-tuning all now on the table, Table 1 shows a practical way to decide between them.

Hyperparameters Explained for Beginners

Hyperparameters are the settings you choose before training starts, as opposed to the parameters the model learns during training. Getting a rough intuition for each one will save you far more time than blindly copying values from a tutorial.

  • Learning rate: how large a step the optimizer takes when updating weights after each batch. Too high causes training to diverge (loss increases or becomes NaN); too low wastes time and may never reach a good solution within your epoch budget. As a rule of thumb: full fine-tuning wants a small learning rate (1e-5 to 5e-5) since you're nudging an already-good model; LoRA tolerates and often needs a larger one (1e-4 to 3e-4) since you're training small adapter matrices from scratch; feature extraction heads can use standard classical ML defaults (whatever your chosen classifier's default is).
  • Batch size: how many examples are processed together before each weight update. Larger batches give more stable gradient estimates and better hardware utilization, but need more memory. If you run out of memory, reduce batch size first, then consider gradient accumulation to simulate a larger effective batch size without the memory cost.
  • Epochs: how many full passes through the training data. Fine-tuning pretrained models typically needs far fewer epochs than training from scratch—often 2-5 for full fine-tuning or LoRA on a moderate dataset. Watch validation metrics, not just training loss, to decide when to stop.
  • Weight decay: a regularization term that discourages weights from growing arbitrarily large, which helps prevent overfitting. A small value like 0.01 is a common, safe default.
  • Warmup steps/ratio: gradually ramping the learning rate up from zero at the start of training, rather than starting at full strength immediately, which improves training stability, particularly for full fine-tuning.
  • Rank (r) and lora_alpha—LoRA-specific, as covered in the section on LoRA.

Gradient Accumulation

If your hardware can't fit the batch size you want, gradient accumulation lets you simulate a larger effective batch by accumulating gradients over several smaller batches before updating weights:

training_args = TrainingArguments(
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    # effective batch size = 4 x 4 = 16
    # ... other arguments
)

This is one of the most useful tricks for anyone training with limited memory—whether that's a laptop's shared unified memory or a modest consumer GPU's VRAM—it lets you train with an effective batch size you couldn't otherwise fit, at the cost of somewhat slower wall-clock training time.

A Practical Hyperparameter Search Strategy for Beginners

You don't need an exhaustive grid search. A pragmatic sequence:

  • Start with the widely used defaults given throughout this article.
  • Train for a small number of epochs (1-2) on a subset of your data to sanity-check that loss is decreasing and nothing is broken.
  • Run the full training with your chosen defaults and observe the training/validation curves.
  • If overfitting, increase weight decay/dropout, reduce epochs, or add more data.
  • If underfitting, increase LoRA rank, unfreeze more layers, increase epochs, or reconsider whether you have enough data for the technique you're using.

Evaluating Your Fine-Tuned Model

A model that “trained without errors” is not the same as a model that actually improved on your task. Evaluation is where you find out which is true.

Standard Metrics

For classification tasks, accuracy, precision, recall, and F1 score (via the evaluate library, as shown in earlier in Part 1 of the series) are the standard starting point. For generative tasks, metrics like BLEU and ROUGE measure surface-level overlap with reference text, but they're known to correlate imperfectly with actual output quality—treat them as one signal among several, not ground truth.

Always Compare Against a Baseline

The single most important evaluation practice, and the one beginners most often skip, is comparing your fine-tuned model's performance against the un-fine-tuned base model on the exact same test set. Without this comparison, you cannot actually claim fine-tuning helped—you can only claim the model produces some output.

Recall the model we trained using LoRA in Part 1, then merged back into the base model with merge_and_unload() and saved to disk as a standalone classifier. Since the task here is predicting a Yelp review's star rating—a 5-class classification problem—the natural metric is accuracy: what fraction of test reviews does each model classify correctly? A short evaluation script is shown in Listing 5.

Listing 5: Evaluating and comparing vase vs. merged LoRA model accuracy on Yelp reviews classification

import torch
from datasets import load_dataset
from sklearn.metrics import accuracy_score
from transformers import (
    AutoModelForSequenceClassification,
    AutoTokenizer,
)
model_name = "distilbert-base-uncased"
merged_model_path = "./my-merged-model"
tokenizer = AutoTokenizer.from_pretrained(
    model_name
)
# Base model, freshly loaded for a fair
# comparison
base_model = AutoModelForSequenceClassification \
  .from_pretrained(
    model_name, num_labels=5
).to(device)
# Your merged LoRA model -- a full
# standalone classifier
merged_model = AutoModelForSequenceClassification \
  .from_pretrained(
    merged_model_path
).to(device)
raw_dataset = load_dataset(
    "Yelp/yelp_review_full"
)
test_dataset = raw_dataset["test"] \
  .shuffle(seed=42).select(range(200))
def evaluate_model(
    model, tokenizer, test_texts, test_labels
):
    model.eval()
    predictions = []
    with torch.no_grad():
        for text in test_texts:
            inputs = tokenizer(
                text,
                truncation=True,
                max_length=256,
                return_tensors="pt",
            ).to(device)
            logits = model(**inputs).logits
            predictions.append(
                torch.argmax(logits, dim=-1).item()
            )
    return accuracy_score(
        test_labels, predictions
    )
base_accuracy = evaluate_model(
    base_model,
    tokenizer,
    test_dataset["text"],
    test_dataset["label"],
)
finetuned_accuracy = evaluate_model(
    merged_model,
    tokenizer,
    test_dataset["text"],
    test_dataset["label"],
)
print(
    f"Base model accuracy:        "
    f"{base_accuracy:.4f}"
)
print(
    f"Fine-tuned model accuracy:  "
    f"{finetuned_accuracy:.4f}"
)
print(
    f"Improvement:                "
    f"{finetuned_accuracy - base_accuracy:+.4f}"
)

For this simple example, I get the following output:

Base model accuracy:        0.2350
Fine-tuned model accuracy:  0.6150
Improvement:                +0.3800

The results are striking: a 38-percentage-point jump in accuracy, from 23.5% to 61.5%. This is exactly the kind of before-and-after comparison that makes fine-tuning gains concrete rather than anecdotal.

Notice that the base model's accuracy—23.5%—is barely better than random guessing across five classes (which would land at 20%). This isn't surprising: distilbert-base-uncased was loaded with a freshly initialized, randomly weighted classification head, so it starts with no notion of how review text maps to star ratings.

After LoRA fine-tuning, the model correctly predicts the star rating for roughly 6 out of every 10 reviews—a substantial and measurable improvement in a task that even humans find genuinely ambiguous at times. A review that reads as “solidly good” could reasonably earn either 3 or 4 stars, and strict accuracy penalizes the model equally whether its prediction is wildly off or just one star away. This is worth keeping in mind when interpreting the number: 61.5% strict accuracy likely understates how useful the model actually is in practice.

Checking for Catastrophic Forgetting

For generative models especially, it's worth checking that your fine-tuned model hasn't lost general capabilities outside your narrow training task. A simple approach: keep a small, fixed set of general-purpose prompts unrelated to your fine-tuning task (basic reasoning, general knowledge, simple instructions) and manually or programmatically compare the base model's and fine-tuned model's responses to those prompts before and after training. A significant, consistent degradation is a sign your learning rate was too high, you trained for too many epochs, or full fine-tuning was too aggressive a choice for your dataset size—LoRA or partial fine-tuning are the natural remedies.

LLM-as-Judge Evaluation for Generative Tasks

For open-ended generative tasks where a single “correct” reference answer doesn't exist, a common and practical technique is to use a strong LLM to score or compare outputs:

def llm_judge_compare(
    prompt, 
    response_a, 
    response_b, 
    judge_client
):
    judge_prompt = f"""\
Compare these two responses to the same
prompt and decide which is better.
Prompt: {prompt}
Response A: {response_a}
Response B: {response_b}
Reply with only "A", "B", or "Tie",
followed by a one-sentence reason."""
    result = 
judge_client.generate(judge_prompt)
    return result

Treat LLM-as-judge results as a useful directional signal, not an infallible ground truth—judge models have their own biases (e.g., a tendency to favor longer responses), so pair this with a smaller sample of human review whenever the stakes justify it.

Model Distillation

So far, we've covered two pieces of the fine-tuning puzzle: adapting a model's behavior through techniques like LoRA and prompt tuning, and then rigorously evaluating whether that adaptation actually helped. Both of these assume the same underlying goal—taking an existing model and making it better at a specific task, whether that's writing in a particular style or classifying reviews by star rating.

Model distillation asks a fundamentally different question. Instead of “how do we make this model better at X,” it asks “how do we make this model smaller while keeping it just as good?” This shift in objective changes almost everything about the training setup—there's no task-specific dataset with ground-truth labels to fit against. Instead, one model teaches another.

Distillation is about compressing knowledge from one model into a different, smaller model. A large, capable teacher model stays completely frozen; a smaller student model is trained to mimic the teacher's output distribution. It's not about improving a task—it's about shrinking a model while preserving as much of its behavior as possible.

The Teacher-Student Idea

Distillation involves two models. The teacher is your large, already capable model, and it stays completely frozen throughout—it is never updated. The student is a smaller, trainable model that learns to mimic the teacher's behavior. Critically, the student doesn't just learn from the raw correct labels the way a model trained from scratch would; it also learns from the teacher's full output distribution, which carries much richer information than a single correct answer does. Figure 2 shows how conceptually model distillation works.

Figure 2: A frozen teacher model and a trainable student model both process the same input, with a distillation loss connecting their outputs
Figure 2: A frozen teacher model and a trainable student model both process the same input, with a distillation loss connecting their outputs

Logits, Probabilities, and Why “Soft” Targets Help

To see why this richer signal matters, it helps to distinguish two things a model produces before its final answer:

  • Logits are the model's raw, unnormalized output scores for each possible answer.
  • Probabilities are what you get after passing logits through a softmax function, which converts them into values that sum to 1.

For example, if a model's logits for the next word are [2.0, 1.0, 0.1] for the candidates “the,” “a,” “an,” softmax converts these into probabilities of roughly 0.66, 0.24, and 0.10. Ordinarily, training only uses the single correct label (a “hard target”—100% “the,” 0% everything else).

Distillation instead trains the student to match the teacher's full probability distribution across all candidates (a “soft target”)—which tells the student not just what the right answer was, but how confident the teacher was, and which wrong answers the teacher considered plausible runners-up.

That's meaningfully more information per training example than a hard label alone provides.

Temperature: Controlling How “Soft” the Targets Are

The softness of these targets is itself tunable, via a temperature parameter T applied inside the softmax before comparing teacher and student outputs. At T = 1, softmax behaves normally and tends to produce a sharp, peaked distribution when the teacher is confident (e.g., logits [10, 2, 0] become roughly [0.9997, 0.00025, 0.00005]—nearly all probability mass on one answer). Raising the temperature (e.g., T = 5) flattens the distribution (the same logits might become roughly [0.88, 0.07, 0.05]), revealing more of the teacher's “opinion” about the relative plausibility of the runner-up answers, which is exactly the extra signal distillation is trying to transfer.

The following code snippet shows how the temperature parameter reshapes a set of logits into probability distributions of varying sharpness, using PyTorch's F.softmax:

import torch
import torch.nn.functional as F
def softened_probabilities(logits, 
    temperature=1.0):
    return F.softmax(logits / temperature,
                     dim=-1)
logits = torch.tensor([10.0, 2.0, 0.0])
print(softened_probabilities(logits, 
      temperature=1))   # sharp
print(softened_probabilities(logits,
      temperature=5))   # softer, more 
                        # informative

Understood.

tensor([9.9962e-01, 3.3533e-04, 4.5383e-05])
tensor([0.7478, 0.1510, 0.1012])

This same temperature parameter, worth noting, is the identical concept behind the “temperature” setting you adjust when sampling text from a chat model at inference time—higher temperature makes output distributions flatter and more varied; lower temperature makes them sharper and more deterministic. Distillation and everyday text generation are using the same underlying mechanism for two different purposes.

Measuring the Loss

A student is typically trained with a combined loss: one term that pulls its output distribution toward the teacher's softened distribution (commonly measured with KL divergence), and another ordinary term that pulls it toward the true ground-truth labels, so the student stays grounded in what's actually correct rather than purely imitating the teacher's occasional mistakes.

Listing 6 shows how to measure the distillation loss between the teacher and student models. The distillation_loss() function implements exactly this combination:

Listing 6: Implementing a knowledge distillation loss function (soft + hard loss)

import torch.nn as nn
import torch.nn.functional as F
def distillation_loss(
    student_logits,
    teacher_logits,
    true_labels,
    temperature=4.0,
    alpha=0.5,
):
    # Soft loss: student learns from
    # the teacher's softened distribution
    soft_teacher = F.softmax(
        teacher_logits / temperature,
        dim=-1,
    )
    soft_student = F.log_softmax(
        student_logits / temperature,
        dim=-1,
    )
    soft_loss = F.kl_div(
        soft_student,
        soft_teacher,
        reduction="batchmean",
    ) * (temperature ** 2)
    # Hard loss: student also learns
    # from the actual ground-truth labels
    hard_loss = F.cross_entropy(
        student_logits, true_labels
    )
    return (
        alpha * soft_loss
        + (1 - alpha) * hard_loss
    )
  • soft_teacher and soft_student divide both models' logits by temperature before applying softmax/log_softmax—this “softens” each distribution, spreading probability mass across more classes so the student can see relative confidence between wrong answers too, not just which single class won.
    F.kl_div then measures how far the student's softened distribution is from the teacher's, and the result is multiplied by temperature ** 2**. That correction matters: dividing logits by T shrinks the gradients flowing from the soft-target term by roughly 1/T², so multiplying the loss back up by T² keeps the soft and hard terms on a comparable scale no matter what temperature is chosen—without it, increasing temperature would silently weaken the teacher's influence rather than just smoothing its distribution.
  • hard_loss, by contrast, uses the original, untempered student_logits against true_labels via ordinary.
  • F.cross_entropy—the student's normal supervised signal, untouched by temperature.
  • The alpha parameter then sets the trade-off between the two: alpha closer to 1 leans more on mimicking the teacher, closer to 0 leans more on the ground truth.

Both models see the same input at every training step; the teacher runs in inference mode only (no gradient updates), while the student's weights are updated by backpropagating this combined loss.

The Distillation Training Loop

With the loss function in hand, the next step is putting it to work in an actual training loop. Rather than fine-tune a teacher from scratch, this example reuses an existing off-the-shelf model as the teacher—nlptown/bert-base-multilingual-uncased-sentiment, which already predicts 1–5 star ratings from review text—and pairs it with a smaller, untrained distilbert-base-uncased student.

The task and dataset are the same yelp_review_full split introduced earlier, but here the goal is different: not steering an existing model's behavior, but compressing the teacher's knowledge into a cheaper model that approximates its accuracy.

One complication is worth calling out before the code: because the teacher and student come from different model families, they don't share a tokenizer or vocabulary, so every training example has to be tokenized twice—once for each model—and both encodings need to travel through the pipeline together. That's handled by a custom tokenization step and a matching custom data collator. The training itself is driven through a DistillationTrainer subclass, which overrides Trainer's compute_loss to run the student forward with gradients, the teacher forward under torch.no_grad(), and combine their logits through distillation_loss from Listing 6 above.

Listing 7 below shows the whole thing end-to-end, from loading the dataset through saving the distilled student.

Listing 7: Distilling a teacher model into a smaller student model

import torch
from datasets import load_dataset
from transformers import (
    AutoTokenizer,
    AutoModelForSequenceClassification,
    Trainer,
    TrainingArguments,
)
def get_device():
    if torch.cuda.is_available():
        # Windows/Linux with NVIDIA GPU
        return torch.device("cuda")
    if torch.backends.mps.is_available():
        # Apple Silicon Mac
        return torch.device("mps")
    # Fallback, works everywhere
    return torch.device("cpu")
device = get_device()
# Full Yelp Review Full dataset -- ~650,000 reviews, 
# 5-star classes
raw_dataset = load_dataset(
    "Yelp/yelp_review_full"
)
# Teacher: an off-the-shelf model, already trained 
# for this task -- no fine-tuning of the teacher
# required. It predicts 1-5 star ratings from review 
# text, matching yelp_review_full's label scheme.
teacher_name = (
    "nlptown/bert-base-multilingual"
    "-uncased-sentiment"
)
teacher_tokenizer = (
    AutoTokenizer.from_pretrained(
        teacher_name
    )
)
teacher_model = (
    AutoModelForSequenceClassification
    .from_pretrained(teacher_name)
    .to(device)
)
teacher_model.eval()
for param in teacher_model.parameters():
    param.requires_grad = False
# Student: a smaller architecture, trained from 
# scratch on this task
student_name = "distilbert-base-uncased"
student_tokenizer = (
    AutoTokenizer.from_pretrained(
        student_name
    )
)
student_model = (
    AutoModelForSequenceClassification
    .from_pretrained(
        student_name, num_labels=5
    ).to(device)
)
# Teacher and student use different vocabularies, 
# so each example is tokenized twice -- once per 
# model -- and both travel through training.
def tokenize_both(examples):
    student_enc = student_tokenizer(
        examples["text"],
        truncation=True,
        max_length=64,
    )
    teacher_enc = teacher_tokenizer(
        examples["text"],
        truncation=True,
        max_length=64,
    )
    return {
        "input_ids":
            student_enc["input_ids"],
        "attention_mask":
            student_enc["attention_mask"],
        "teacher_input_ids":
            teacher_enc["input_ids"],
        "teacher_attention_mask":
            teacher_enc["attention_mask"],
        "labels": examples["label"],
    }
tokenized_train = raw_dataset["train"].map(
    tokenize_both,
    batched=True,
    remove_columns=["text", "label"],
)
def dual_collator(features):
    def pad(ids_key, mask_key):
        max_len = max(
            len(f[ids_key])
            for f in features
        )
        input_ids, attn = [], []
        for f in features:
            pad_len = (
                max_len - len(f[ids_key])
            )
            input_ids.append(
                f[ids_key] + [0] * pad_len
            )
            attn.append(
                f[mask_key] + [0] * pad_len
            )
        return (
            torch.tensor(input_ids),
            torch.tensor(attn),
        )
    student_ids, student_mask = pad(
        "input_ids", "attention_mask"
    )
    teacher_ids, teacher_mask = pad(
        "teacher_input_ids",
        "teacher_attention_mask",
    )
    labels = torch.tensor(
        [f["labels"] for f in features]
    )
    return {
        "input_ids": student_ids,
        "attention_mask": student_mask,
        "teacher_input_ids": teacher_ids,
        "teacher_attention_mask":
            teacher_mask,
        "labels": labels,
    }
class DistillationTrainer(Trainer):
    def __init__(
        self, *args, teacher_model=None,
        temperature=4.0, alpha=0.5,
        **kwargs
    ):
        super().__init__(*args, **kwargs)
        self.teacher_model = teacher_model
        self.temperature = temperature
        self.alpha = alpha
    def compute_loss(
        self, model, inputs,
        return_outputs=False, **kwargs
    ):
        labels = inputs["labels"]
        student_outputs = model(
            input_ids=inputs["input_ids"],
            attention_mask=
                inputs["attention_mask"],
        )
        student_logits = (
            student_outputs.logits
        )
    
        with torch.no_grad():
            teacher_logits = (
                self.teacher_model(
                    input_ids=inputs[
                        "teacher_input_ids"
                    ],
                    attention_mask=inputs[
                        "teacher_attention"
                        "_mask"
                    ],
                ).logits
            )
    
        # distillation_loss
        loss = distillation_loss(
            student_logits,
            teacher_logits,
            labels,
            temperature=self.temperature,
            alpha=self.alpha,
        )
        return (
            (loss, student_outputs)
            if return_outputs else loss
        )
training_args = TrainingArguments(
    output_dir="./distilled-student",
    per_device_train_batch_size=16,
    num_train_epochs=1,
    learning_rate=5e-5,
    logging_steps=200,
    report_to="none",
    remove_unused_columns=False,
    # keep teacher_input_ids and
    # teacher_attention_mask
)
trainer = DistillationTrainer(
    model=student_model,
    args=training_args,
    train_dataset=tokenized_train,
    data_collator=dual_collator,
    teacher_model=teacher_model,
    temperature=4.0,
    alpha=0.5,
)
trainer.train()
student_model.save_pretrained(
    "./distilled-student-final"
)
student_tokenizer.save_pretrained(
    "./distilled-student-final"
)

Understood.

  • get_device()—the same CUDA/MPS/CPU fallback helper used throughout the article, so the code runs unmodified on an NVIDIA GPU, an Apple Silicon Mac, or plain CPU.
  • Dataset—loads the full Yelp/yelp_review_full split (~650,000 labeled reviews, 5-star classes).
  • Teacher—nlptown/bert-base-multilingual-uncased-sentiment, a public model already trained to predict 1–5 star ratings from text. It's frozen two ways: eval() mode (no dropout/batchnorm drift) and requires_grad = False on every parameter (no gradient flow), so it never gets updated during training—it only ever supplies soft targets.
  • Student—a fresh, untrained distilbert-base-uncased classification head with five output labels, the model actually being trained.
  • tokenize_both—because the teacher and student come from different model families, they don't share a vocabulary or tokenizer. Every example gets encoded twice—once per model—and both encodings are carried through the dataset under separate column names (teacher_input_ids / teacher_attention_mask vs. the student's default input_ids / attention_mask).
  • dual_collator—a custom batch collator, since the built-in DataCollatorWithPadding only knows how to pad one set of input_ids. This one pads both the student's and teacher's token sequences independently, plus batches the labels.
  • DistillationTrainer—subclasses Trainer and overrides compute_loss: it runs the student forward (with gradients), the teacher forward (inside torch.no_grad(), since it never learns), then combines their logits with the ground-truth labels via the distillation_loss function from earlier—blending “learn from the teacher's soft distribution” and “learn from the actual correct answer” into one loss.
  • remove_unused_columns=False—the fix from your last run: by default, Trainer strips any dataset column that isn't a named argument of the model's forward() method, which would silently delete teacher_input_ids/teacher_attention_mask before the collator ever sees them. This flag disables that stripping.
  • Training and saving—trains for one epoch over the full 650K-review dataset, then saves the distilled student model and its tokenizer to disk—a smaller, cheaper model that's learned to approximate the larger off-the-shelf teacher's behavior on this specific task.

Exporting and Deploying Your Model

Once you have a fine-tuned model you're happy with, deployment options span from “load it directly in Python” to fully optimized, cross-platform local inference. I will outline two options below:

  • Direct transformers Inference
  • Exporting to ONNX for Cross-Platform Inference

The Simplest Path: Direct Transformers Inference

For many applications, especially internal tools and prototypes, simply loading the saved model directly is sufficient and requires no additional export step:

import torch
from transformers import pipeline
def get_device():
    if torch.cuda.is_available():
        # Windows/Linux with NVIDIA GPU
        return torch.device("cuda") 
    if torch.backends.mps.is_available():
        # Apple Silicon Mac
        return torch.device("mps")
    # Fallback, works everywhere
    return torch.device("cpu") 
classifier = pipeline(
    "text-classification",
    model="./my-merged-model",
    tokenizer="./my-merged-model",
    device=get_device(),
)
result = classifier(
    "This product exceeded my expectations!"
)
print(result)

For the above code snippet to work, you must have a folder named my-merged-model in the same folder as your current code. The my-merged-model is a model that we trained in Part 1 of this article series, under the “Merging Adapters for Deployment” section.

Exporting to ONNX for Cross-Platform Inference

For production deployment where you want fast inference without a Python/PyTorch dependency, or need to run the same exported model across Windows, Mac, and Linux consistently, ONNX (Open Neural Network Exchange) is a mature, platform-agnostic option.

ONNX Runtime automatically selects an appropriate execution provider for whatever hardware it's running on (CPU, CUDA, DirectML on Windows, CoreML on Mac), which makes it a genuinely platform-agnostic deployment target once exported.

To export your model to the ONNX format, you first need to install the optimum-onnx package:

pip install "optimum-onnx[onnxruntime]"

Then, in Terminal, use the following command to export the “my-merged-model” to ONNX format:

optimum-cli export onnx --model \
 ./my-merged-model ./onnx-model \
 --task text-classification

The converted model would now be saved in the “onnx-model” folder. Listing 8 shows how you can load this exported ONNX model back and run inference with it using Hugging Face's pipeline API, exactly as you would with a regular PyTorch model.

Listing 8: Loading and running inference with the exported ONNX model

from optimum.onnxruntime import (
    ORTModelForSequenceClassification,
)
from transformers import (
    AutoTokenizer,
    pipeline,
)
onnx_model = (
    ORTModelForSequenceClassification
    .from_pretrained("./onnx-model")
)
tokenizer = AutoTokenizer.from_pretrained(
    "./onnx-model"
)
onnx_classifier = pipeline(
    "text-classification",
    model=onnx_model,
    tokenizer=tokenizer,
)
print(
    onnx_classifier(
        "This is a fantastic result."
    )
)

Understood.

[
  {
    'label': 'LABEL_4', 
    'score': 0.7311358451843262
  }
]

Note that for this example to work, you need to save the tokenizer after the merge step (under the “Merging Adapters for Deployment” section in Part 1 of this article series):

merged_model = model.merge_and_unload()
merged_model.save_pretrained(
    "./my-merged-model")
# add this to save the tokenizer
tokenizer.save_pretrained("./my-merged-model")

The ONNX export process (via optimum-cli export onnx) doesn't just convert the model's weights—it also copies over the tokenizer files from the source directory into the output ONNX folder, so that both can be loaded together later with a single, consistent set of files. If tokenizer.save_pretrained("./my-merged-model") was never called, the ./my-merged-model directory would contain only the model weights and config (from merged_model.save_pretrained(...))—no tokenizer.json, vocab.txt, or tokenizer_config.json. Since the exporter reads directly from that directory, it would have no tokenizer files to copy, and the resulting ./onnx-model folder would be missing them too.

Choosing a Deployment Path

Table 2 summarizes the two deployment targets covered, and when to reach for each.

Beyond Text: Feature Extraction and Fine-Tuning for Images

Everything in this article has focused on text, but the same four-technique spectrum—feature extraction, partial fine-tuning, full fine-tuning, and LoRA—applies just as directly to vision models, and it's worth walking through briefly, both because image classification is a common real-world need and because seeing the identical pattern in a different domain reinforces just how general the underlying idea is.

A simple analogy. Imagine a photo as a big LEGO puzzle. A vision transformer first cuts the picture into small squares—little LEGO pieces—and turns each square into a number code (a feature vector) describing what's in it: a color, an edge, part of a shape. It then looks at all the squares together and figures out how they relate—realizing, say, that a square with an ear and a square with eyes belong to the same dog—and combines that information into a final guess about the whole picture. Feature extraction, in this picture, is: cut into squares → turn into number codes → look at all of them together → make a decision—exactly the same “frozen backbone, trainable head” pattern from earlier sections, just applied to patches of pixels instead of tokens of text.

Feature Extraction with a Vision Transformer

Just as we extracted sentence embeddings from a frozen text backbone in Part 1 of this article, we can extract image embeddings from a frozen vision transformer (ViT), as shown in Listing 9.

Listing 9: Extracting image embeddings with a pretrained vision transformer (ViT)

from transformers import (
    AutoImageProcessor,
    AutoModel,
)
from PIL import Image
import torch
def get_device():
    if torch.cuda.is_available():
        # Windows/Linux w/ NVIDIA GPU
        return torch.device("cuda")
    if torch.backends.mps.is_available():
        # Apple Silicon Mac
        return torch.device("mps")
    # Fallback, works everywhere
    return torch.device("cpu")
device = get_device()
model_name = (
    "google/vit-base-patch16-224"
)
processor = (
    AutoImageProcessor.from_pretrained(
        model_name
    )
)
vit_model = (
    AutoModel.from_pretrained(
        model_name
    ).to(device)
)
vit_model.eval()
def extract_image_embedding(
    image_path,
):
    image = Image.open(
        image_path
    ).convert("RGB")
    inputs = processor(
        images=image,
        return_tensors="pt",
    ).to(device)
    with torch.no_grad():
        outputs = vit_model(**inputs)
    # The [CLS] token's final hidden
    # state is a strong whole-image
    # summary vector
    cls_embedding = (
        outputs.last_hidden_state[
            :, 0, :
        ]
    )
    return (
        cls_embedding.cpu()
        .numpy()
        .squeeze()
    )
embedding = extract_image_embedding(
    "example.jpg"
)
print(embedding.shape)  # (768,)
print(embedding)        # vector of 768 numbers  

Understood.

From here, you can extract embeddings for your full labeled image dataset, then train a logistic regression or small neural network head on top. This is an excellent starting point for tasks like product categorization, document classification by appearance, or defect detection, especially when you have a modest number of labeled images per class.

Fine-Tuning a Vision Model with LoRA

When feature extraction isn't enough—for example, fine-grained classification where subtle visual differences matter—LoRA applies to vision transformers exactly as it does to text models, targeting the attention projection layers, as shown in Listing 10.

Listing 10: Fine-tuning a vision model with LoRA

from transformers import (
    AutoModelForImageClassification,
)
from peft import LoraConfig, get_peft_model
model = \
    AutoModelForImageClassification.from_pretrained(
      "google/vit-base-patch16-224",
      num_labels=10,
      # replaces the 1000-class head with
      # your 10-class head
      ignore_mismatched_sizes=True,
)
lora_config = LoraConfig(
    r=8,
    lora_alpha=16,
    lora_dropout=0.1,
    target_modules=["query", "value"],
    modules_to_save=["classifier"],
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()

Training then proceeds through the same Trainer API used throughout this article, with an image-specific data collator handling batching of pixel tensors instead of token IDs. The conceptual throughline—frozen backbone, small trainable adapters, dramatically reduced memory footprint compared to full fine-tuning—carries over completely unchanged from text to vision.

Multimodal Models: CLIP as an Example

Models like CLIP, which jointly embed images and text into a shared representation space, extend the feature-extraction pattern to cross-modal tasks—image-text retrieval, zero-shot image classification via text prompts, and similarity search, as shown in Listing 11.

Listing 11: Zero-shot image classification with CLIP

from transformers import CLIPModel, CLIPProcessor
clip_model = CLIPModel.from_pretrained(
    "openai/clip-vit-base-patch32"
).to(device)
clip_processor = CLIPProcessor.from_pretrained(
    "openai/clip-vit-base-patch32"
)
def get_image_and_text_embeddings(
    image, candidate_labels
):
    inputs = clip_processor(
        text=candidate_labels,
        images=image,
        return_tensors="pt",
        padding=True,
    ).to(device)
    with torch.no_grad():
        outputs = clip_model(**inputs)
    # Cosine similarity between image
    # and each text label's embedding
    logits_per_image = (
        outputs.logits_per_image
    )
    probs = logits_per_image.softmax(
        dim=1
    )
    return probs
probs = get_image_and_text_embeddings(
    Image.open("example.jpg"),
    [
        "a photo of a cat",
        "a photo of a dog",
        "a photo of a car",
    ],
)
print(probs)

CLIP (Contrastive Language-Image Pre-training) is a neural network model developed by OpenAI that learns to connect images and text by training on a large dataset of image-caption pairs found on the internet.

In the above code listing, passing in an example photo of a dog results in the following:

tensor([
  [3.2746e-04, 9.9964e-01, 2.9310e-05]
], device='mps:0')

The result is correct as the “a photo of a dog” yields the highest probability of 9.9964e-01.

This “zero-shot” use of CLIP—classifying images against arbitrary text labels with no task-specific training at all—is worth trying before any fine-tuning effort on image classification tasks, for exactly the same reason prompting is worth trying before fine-tuning text models: it's free, instant, and often good enough. Fine-tuning CLIP with LoRA is the natural next step when zero-shot performance falls short on your specific label set.

Glossary

This article uses a number of terms—LoRA, quantization, catastrophic forgetting, and others—with specific technical meanings that don't always match their everyday usage. To keep the main text focused, I've collected precise definitions here rather than re-explaining each term inline every time it comes up. If you hit an unfamiliar word while reading, or just want a quick refresher, the glossary below covers every key term used in these two articles.

  • Adapter—In LoRA and related PEFT methods, the small set of trainable matrices injected into a frozen model. “Adapter” and “LoRA adapter” are used interchangeably in this article.
  • Backbone—The main body of a transformer model (its stacked attention/feed-forward layers), as distinct from a task-specific head attached on top.
  • Catastrophic forgetting—The degradation of a model's general capabilities that can occur when aggressive fine-tuning overwrites broadly useful pretrained representations with narrow, task-specific ones.
  • Checkpoint—A saved snapshot of a model's weights at a particular point, either from pretraining or fine-tuning.
  • Embedding—A fixed-length numerical vector representation of an input (a word, sentence, or image) produced by a model, capturing semantic information in a form usable by downstream algorithms.
  • Feature extraction—Using a frozen pretrained model purely to produce embeddings, then training a separate lightweight model on those embeddings, without updating the pretrained model itself.
  • Fine-tuning—Continuing to train a pretrained model's parameters (fully or partially) on new, typically smaller and more task-specific, data.
  • Gradient accumulation—Simulating a larger training batch size by summing gradients over several smaller batches before performing a weight update, used to work around memory constraints.
  • Head—A small, typically randomly-initialized final layer (or few layers) mapping a model's internal representations to a specific output, such as class probabilities.
  • LoRA (Low-Rank Adaptation)—A parameter-efficient fine-tuning technique that freezes the pretrained model and learns small low-rank matrices approximating the weight updates full fine-tuning would otherwise apply directly.
  • Overfitting—When a model learns patterns specific to its training data that don't generalize to new data, visible as training performance improving while validation performance stagnates or worsens.
  • Quantization—Reducing the numerical precision used to store model weights (e.g., from 16-bit to 4-bit), trading a small amount of accuracy for large reductions in memory footprint.
  • QLoRA—The combination of quantized model loading with LoRA fine-tuning, enabling fine-tuning of large models on modest hardware.
  • Rank (r)—In LoRA, the dimensionality of the low-rank adapter matrices; a key trade-off between adapter capacity and parameter efficiency.
  • Tokenizer—The component that converts raw text into the numerical token IDs a model consumes, and must always match the model it's paired with.

Summary

Fine-tuning is not one technique but a spectrum, and the biggest lever most developers can pull isn't a hyperparameter—it's choosing the right point on that spectrum for their task, their data, and their hardware. Feature extraction, partial fine-tuning, full fine-tuning, and LoRA all have a legitimate place; the practical skill is knowing which one to reach for first, and having a fast, honest way to tell whether it worked.

The workflow this series of articles has walked through—start with feature extraction, escalate to LoRA if needed, reserve full fine-tuning for when you've measured that it's genuinely justified, and always evaluate against a real baseline—scales from a weekend side project to a production system, on whatever hardware you happen to have: a Windows gaming PC, a Mac on your desk, or a Linux workstation.

Understood.

Distillation Vs. the Other Techniques in This Article

It's worth being clear about how distillation relates to everything else covered so far: feature extraction, partial/full fine-tuning, LoRA, and prompt tuning are all about adapting an existing model to a new task.

Distillation is orthogonal to that—it's about compressing an existing (possibly already fine-tuned) model into a smaller, cheaper one that preserves as much of its behavior as possible.

In practice, the two are often combined: fine-tune a large model to be excellent at your task using LoRA or full fine-tuning, then distil that fine-tuned teacher down into a small student model that's cheap enough to serve at scale or run on-device, accepting a modest accuracy trade-off in exchange for a significantly smaller, faster model.

Understood.

<a id="table1"Table 1: Deciding between the different ways of PEFT-style options

Choose…

Understood.

Prompt tuning You need maximum parameter efficiency; you're prototyping rapidly; compute is very limited; you need to deploy many task-specific behaviors cheaply; moderate performance is acceptable

Understood.

LoRA

You need meaningfully better task performance than prompt tuning offers; you can afford a modest increase in trainable parameters and compute; the task benefits from deeper adaptation than steering the input alone can provide

Full fine-tuning

You need the maximum possible performance on a specific, often highly specialized domain (medical, legal, scientific); you have abundant data and compute; the task is different enough from the base model's training that surface-level adaptation isn't sufficient

**** Please provide the Original and Marked text to review.

Format:
Original: …
Marked: …

Deployment target Best for Platform behavior
Prototypes, internal tools, Python-based services Identical code on all platforms; picks up CUDA/MPS/CPU automatically

Understood.
Understood.

ONNX Runtime Production services needing fast, dependency-light inference Same exported model runs everywhere; execution provider adapts to hardware
Understood.