> Home/Case Studies/Local AI Platform
Technical Case Study — XponentShift Limited

Local AI Platform

Fine-tuning and retrieval-augmented generation, engineered in-house and run entirely on our own hardware — no cloud model, no per-token bill, no document ever leaving the building.

Fine-Tuning (QLoRA)Retrieval-Augmented GenerationQwen2.5-3B-InstructpgvectorOllamaNext.jsDocker + NVIDIA GPUCloudflare Tunnel
100%
Local Inference
81.7%
Fine-Tune Pass Rate
2%
Fabrication Rate
2,312
Training Examples
The Platform

Two Ways We Ship AI

We run two AI systems entirely on our own infrastructure. One fine-tunes a small open model on a client's own product data until it answers support questions correctly. The other grounds a model in a specific document — a financial report, a filing — so it answers only from what's actually on the page.

Both run on a single consumer GPU, with nothing sent to a third-party model provider.

✗ The Off-The-Shelf Way

Wire up a general-purpose cloud model, prompt it with a system message, and hope it doesn't invent a feature, a page or a policy that doesn't exist — while every question leaves the building and shows up on a bill.

✓ Our Way

Fine-tune a 3-billion-parameter open model on the product's real data, or ground it in the real document, so an answer is either learned from the truth or read directly off it. Inference runs on hardware we own.

▣ Built By

XponentShift Limited — the same in-house engineering team that builds our client products builds the AI layer that runs inside them: dataset generation, fine-tuning, retrieval engineering and evaluation, end to end, with no outsourced model vendor in the loop.

Engineering Edge

How This Differs From
a Bolted-On Chatbot

Typical AI Integration
Our Approach
✗Generic model, answers from public training data
✓Fine-tuned on 2,312 examples generated straight from the product's own code, database schema and copy
✗Cloud API — per-token billing, every question leaves the network
✓100% local inference on owned hardware — zero API cost, zero data egress
✗Confident answers whether or not they're true
✓Every number and fact must appear in the retrieved context; the model is instructed to say "I don't know" instead of guessing
✗Single-shot question and answer
✓Multi-turn memory that automatically rewrites bare follow-ups ("and net income?") into a question the retrieval layer can actually use
✗Wide financial tables read column-blind — nearest header wins
✓Every table is rewritten so each figure is bound to its own explicit reporting period before the model ever sees it
✗"Ship it and see" fine-tuning
✓Baseline measured before training, every claim of improvement graded by machine-checkable rules, exported to a human-reviewable workbook
Engineering Breakdown

What We
Actually Built

Six pieces of engineering, each solving a specific, measured failure — not generic AI-integration boilerplate.

Dataset Engineering

2,312 synthetic training examples generated programmatically from the live product, its database schema, and its real navigation, pricing and status values — never from generic knowledge of how the product works. An independent validator rejects the build on any contradiction of a pinned product fact, any cross-split data leakage, or any invented navigation path.

Fine-Tuning Pipeline

Supervised fine-tuning on Qwen2.5-3B-Instruct, trained on a free Google Colab T4 GPU end to end, merged and exported to GGUF via llama.cpp, and imported into Ollama behind an explicit chat template — small enough to run on a laptop GPU, tuned enough to know the product cold.

Document Ingestion

A custom pdfplumber pipeline extracts prose and tables page by page, repairing rendering artifacts most pipelines miss entirely: double-struck bold text, letter-spaced section headings with no ruling line, and decorative panels that parse as "tables" but hold no data.

Retrieval Engineering

bge-m3 embeddings (1024-dimension, 8,192-token context) stored in Postgres via pgvector, so a dense financial table is one vector instead of being split mid-grid. Retrieval blends cosine similarity with lexical term overlap and reserves dedicated slots for tables.

Conversational Memory

"And is that a monthly charge?" is rewritten into a self-contained question before retrieval ever runs, with deterministic safety checks that reject any rewrite inventing a period, company or figure nobody in the conversation actually said.

Evaluation & Infrastructure

A 260-question held-out benchmark plus a separate 52-question "real feature" set, both rule-graded and exported to a reviewable Excel workbook. Deployed on Docker with NVIDIA GPU passthrough and exposed for stakeholder review over a Cloudflare Tunnel — no cloud hosting bill required.

Engineering Challenges

Four Problems That Don't
Show Up in a Demo

01

Reading the Wrong Column

A financial table spanning five quarters put the right line item in front of the model — with the wrong period's figure sitting right next to it. Asked for Q2-2026 Adjusted EBITDA, it returned the year-ago number because both columns shared the label "Q2," a pattern traced to 15 of 44 evaluation errors. The fix rewrites every table so each figure states its own line item and period inline — "Adjusted EBITDA: Q2-2026 = 3,273" — before the model ever reads it, removing the column arithmetic entirely.

02

Follow-Up Questions With No Noun in Them

"And net income?" carries no company, no period and no subject of its own — embedded and searched literally, it retrieves nothing useful and can route an entire conversation to the wrong document. The fix rewrites every dependent-sounding message into a standalone question using the conversation's own established subject and period, then checks it against two hard rules: nothing in the rewrite may be a number nobody said, and nothing the user asked about may be dropped.

03

A Benchmark That Rewarded Saying No

A third of the held-out evaluation set scored the model on correctly withholding information, so a fine-tune that simply got more cautious scored higher on paper — while a real support conversation, full of ordinary product questions, is exactly the slice that benchmark barely tested. The fix built a second evaluation set of real, answerable product questions specifically to catch a model that scores well by refusing more; a release decision reads both sets together, never either alone.

04

Four Gigabytes of GPU Memory

A side-by-side base-vs-fine-tuned comparison tool needs two 3-billion-parameter models loaded at once — which don't both fit in VRAM on a consumer GPU. The fix pins Ollama to one loaded model at a time; the second request queues and waits for the swap rather than crashing the container, so the whole platform, including a live model comparison, runs on a single laptop GPU.

Process

From Raw Data to a
Deployed, Evaluated Model

1

Baseline

Measure how the untuned base model performs on the actual target task before writing a single training example — so an improvement claim is measured against a real number, not assumed.

2

Dataset

Generate a synthetic supervised fine-tuning dataset straight from the product's own code, schema and copy. Validate it for contradictions, cross-split leakage and any real personal data before it touches a training run.

3

Train

QLoRA fine-tuning on a free-tier GPU, checkpointed regularly and synced off the training VM as it runs, so an interrupted session resumes instead of losing hours of compute.

4

Evaluate

Score the fine-tune against the base model on a held-out set using machine-checkable rules, plus a second set built to catch over-caution. Export every answer to a workbook a human can actually read.

5

Deploy

Import the trained model into Ollama behind an explicit chat template, serve it through our Next.js stack on Docker with GPU passthrough, and expose it for review over a Cloudflare Tunnel — no cloud GPU bill.

Roadmap

Where This Goes Next

Near-Term
  • Second training round covering prompt-injection handling and stale-state awareness
  • Onboarding-flow coverage, the largest remaining gap in real product questions
  • Regrade tooling so a grading-rule fix never requires re-running the model
Mid-Term
  • Multi-document, multi-company retrieval at scale, beyond a single ingested report
  • Automated regression evaluation on every dataset change, not just before a release
  • A larger local hardware tier for concurrent model serving
Long-Term
  • Fine-tuning-as-a-service: the same pipeline applied to any client's own product data
  • Packaged on-prem deployments for clients with strict data-residency or GDPR requirements
  • Routing between a small tuned model and a larger reasoning model by query difficulty

// Where fine-tuning meets retrieval — and neither one leaves the building

Want AI like this
built for your business?

Designed and built entirely by XponentShift Limited — the same in-house team, the same standard of measurement, applied to your product's own data.