All posts

Behavioral Fine-Tuning: From Raw Correspondence to Deployable Persona in Six Minutes

Dr. Jerry A. Smith · February 25, 2026 · 6 min read

Building Minds — Technical Edition

This is the technical companion to Edition 8, "The Training Corpus Is the Product." That piece explained why behavioral fine-tuning matters for encoding institutional knowledge. This one describes how — the complete pipeline, from raw communication artifacts to a deployable model that runs on a laptop.

The published model is on HuggingFace. Everything described below is reproducible on any Mac with Apple Silicon.

The Problem with Prompted Personas

The standard approach to making a language model behave like a specific person is to describe that person in a system prompt. This works the same way describing a character to an actor does — it produces a performance. The model follows instructions about a persona. It does not become the persona.

The limitation is structural. A system prompt operates at inference time. It constrains the model's output distribution through soft guidance, but the underlying weights — the model's actual learned behaviors — remain unchanged. Every response is a negotiation between what the prompt requests and what the weights prefer. Adversarial inputs can override the prompt. Long conversations dilute it. The persona is a mask, not a face.

Behavioral fine-tuning changes the weights. The model's default behavior — the thing it does when no prompt constrains it — becomes the target persona. The distinction matters in the same way that the difference between a musician reading sheet music and a musician playing from memory matters. Both produce notes. One of them can improvise.

Architecture

The pipeline has five stages. Each produces an artifact that feeds the next.

Raw Corpus → Behavioral Shaping → ChatML Formatting → QLoRA Training → GGUF Export

Total wall-clock time from raw data to deployable model: approximately six minutes on an M-series MacBook. The bottleneck is the GGUF dequantization step, not the training.

Stage 1: Data Pipeline

Ingestion

The corpus ingester accepts .eml, .mbox, .txt, .json, and .jsonl files. It applies zero text normalization. This is a deliberate design choice, not an oversight.

The conventional instinct is to clean training data — correct spelling, normalize punctuation, and remove formatting artifacts. For behavioral fine-tuning, this destroys the corpus's most valuable signal. A person's typos carry information about their cognitive processing speed under load. Their abbreviations reveal social register. Their punctuation patterns encode rhetorical style. Research presented at NAACL 2025 measured the cost: models trained on cleaned correspondence achieve a stylometric F1 of 0.40. The same data, uncleaned, produces 0.53. A thirty-two percent improvement in persona fidelity from doing less work.

The ingester's output is the raw text, exactly as written.

Behavioral Shaping

Every communication sample carries an implicit behavioral posture. Some deflect. Some assert. M

ost are neutral. The behavioral shaper classifies each sample using lightweight signal detection — keyword patterns associated with deflection, assertion, and neutral communication — then resamples the corpus to match target ratios.

The default composition:

These ratios are not arbitrary. They produce a model whose default behavioral posture mirrors the distribution of postures in the training data. If the source person deflects in 30% of their communications, the model deflects in roughly 30% of its responses. The personality is not programmed. It is statistically induced.

For a real corpus, you would measure the natural distribution first and decide whether to preserve it or adjust it. The shaper handles both cases.

ChatML Formatting

All training data is converted to Qwen3's native ChatML template — system, user, assistant roles — and split into train (90%), validation (5%), and test (5%) sets. Samples that include information-seeking patterns are formatted as multi-turn conversations with tool-call structure, teaching the model to retrieve information while maintaining persona voice.

Stage 2: Training

Why QLoRA on MLX

Full fine-tuning of a four-billion-parameter model requires approximately thirty-two gigabytes of memory and hours of compute. QLoRA — quantized low-rank adaptation — achieves ninety percent of full fine-tuning quality using a fraction of the resources. The base model is loaded in four-bit precision. Small trainable adapter matrices are inserted alongside the frozen weights. Only the adapters are trained. The base model never changes.

MLX is Apple's machine learning framework designed specifically for Apple Silicon's unified memory architecture. There is no CPU-to-GPU transfer overhead because there is no separate GPU memory. The model and the training data occupy the same physical memory space. Training runs natively on Metal GPU cores.

Configuration

base_model: Qwen/Qwen3-4B (4.02B parameters)
quantization: 4-bit (mlx-community/Qwen3-4B-4bit)
adapter: LoRA, rank 16, alpha 32
target_modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
layers: 8 (top layers, where linguistic style is encoded)
learning_rate: 2e-4, cosine schedule
batch_size: 2, gradient accumulation 8 (effective batch 16)
epochs: 3

Rank 16 is the empirical sweet spot for persona tasks. Rank 8 underfits stylistic patterns — the model captures vocabulary but not rhythm. Rank 32 overfits on small corpora — the model reproduces training examples rather than generalizing the voice. The eight target layers are the top of the network, where high-level linguistic features are represented. Lower layers encode syntax and grammar that should remain unchanged.

Results

The gap between training and validation loss is minimal — 0.079 versus 0.070 — indicating generalization without overfitting. Convergence is rapid: most learning occurs in the first epoch. Epochs two and three refine the adaptation.

Stage 3: Evaluation

Stylometric F1

The primary quantitative metric. Twelve linguistic features are extracted from both the source corpus and the model's generated output:

Lexical: average word length, average sentence length, vocabulary richness. Punctuation: comma rate, em-dash frequency, question ratio, exclamation ratio, ellipsis rate. Structural: average paragraph length, sentence length variance, and short sentence ratio. N-gram: bigram distribution overlap.

The composite F1 combines cosine similarity between feature vectors (60%), per-feature deviation scoring (25%), and bigram Jaccard overlap (15%).

The F1 of 0.616 exceeds both the target and the NAACL 2025 baseline of 0.53 for raw-fidelity training. The low bigram Jaccard is correct behavior — the model generates novel phrasing in the persona's style rather than reproducing memorized sequences.

Big Five Personality Stability

Twenty trait-probing prompts mapped to OCEAN dimensions. Cronbach's alpha measures whether the model exhibits consistent personality traits across diverse contexts. The current result — α = 0.200 — is below the 0.70 target, expected with synthetic training data. Real correspondence, with its natural consistency of personality across thousands of genuine interactions, produces substantially higher stability.

Style Drift During Tool Use

When a model retrieves external information mid-conversation, persona fidelity typically degrades. The response reverts toward the generic voice of the base model. Style drift is measured as the stylometric difference between tool-augmented and non-augmented responses. Target: less than ten percent deviation.

Stage 4: Deployment

Adapter Fusion

The LoRA adapters — 14 megabytes of trained weights — are mathematically combined with the base model. The result is a standalone model that requires no adapter loading at inference time. This is a one-way operation: the merged model cannot be separated back into base and adapter.

GGUF Export

The fused model is dequantized to BFloat16 precision, converted to GGUF format via llama.cpp's conversion script, then quantized to the target precision.

Q8_0 is recommended for behavioral models because the subtle stylistic patterns that define persona fidelity — punctuation rhythms, sentence length distributions, vocabulary selection — are the first features that aggressive quantization blurs.

The resulting file runs in LM Studio, Ollama, or any llama.cpp-compatible runtime. No cloud. No API. No data leaves the machine.

The Pipeline in Five Commands

pip install -r requirements.txt
python scripts/generate_sample_data.py --count 900
python scripts/train.py
python scripts/evaluate.py
python scripts/export_gguf.py --quantization q8_0

Replace generate_sample_data.py with real corpus data in data/raw/ for production use. The minimum viable corpus is five hundred samples. Five thousand is more than sufficient. The constraint is authenticity, not volume.

Adapting This to Your Organization

The pipeline was built as a proof-of-concept using synthetic data. Moving to production requires one change: real data.

The data that matters is not the polished, official communications that organizations typically archive. It is the working communications — the emails written quickly, the Slack messages sent without editing, the meeting transcripts where someone thought out loud. These are the artifacts that carry a behavioral signal. They are, in most organizations, the data that no one thought to preserve because no one recognized it as valuable.

The firms that begin collecting this data now — not cleaning it, not normalizing it, simply preserving the raw working communications of their most effective practitioners — will have the training corpus when they are ready to build. The firms that wait will find that the data they need has been overwritten, archived to inaccessibility, or simply never saved.

The training corpus is the product. The pipeline is the packaging. The data collection decision cannot be made retroactively.


Model

Dr. Jerry A. Smith builds AI organizations within large enterprises. Connect on LinkedIn or reach out at jerry@drjerryasmith.com.

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call