Build Your Own Jev With Claude Opus 5.5
Mark Kashef · 2026-09-23 · tutorial · 42,895 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Mark Kashef demonstrates how to build a local, open-source multimodal classifier pipeline inspired by Jev using Claude Opus 5.5 and open-source models. He details an end-to-end workflow to fine-tune an encoder model (such as ModernBERT) to evaluate travel terms, verify photo evidence, and match client requirements locally.
What is shown
- [00:00 - 00:35] Demo of "Away Together," a travel agency app matching 12 customer profiles against hotel packages and cancellation terms.
- [01:02 - 02:08] Breakdown of classification queries (cancellation refund, late arrival, pool access, wheelchair accessibility) and the 4-step framework.
- [02:52 - 04:15] Whiteboard explanation of encoder-only vs. decoder-only architectures and context priming.
- [04:16 - 04:48] Open-source model alternatives shown on Hugging Face and GitHub, including
ModernBERT-base-zeroshot-v2.0and Diffusion Gemma. - [05:34 - 07:04] Prompts and instructions provided to Claude to configure local training, evaluation benchmarks, and image recognition.
- [07:05 - 09:54] The 8-part prompt structure (Job, Computer, Data, Baseline, Training, Final test, App + Images, Delivery) for Claude.
- [09:55 - 10:55] Visual diagram explaining overfitting risk and separating test/validation sets.
- [11:04 - 11:49] JSON data format structure with classification criteria (
meets,violates,insufficient_evidence). - [12:08 - 12:43] Accuracy comparison charts: first model (60.28%), V2 model (95.28%), and closed Jev model (98.61%).
- [12:44 - 13:26] Image verification flow overriding text classification (e.g., detecting steps or identifying a pond instead of a pool).
- [13:27 - 14:22] Querying SuperGrok to locate recent open-source Jev derivatives on GitHub and generating an automated training system prompt for Claude Opus 5.5.
Claims & numbers
- The presenter claims the system runs entirely locally on consumer hardware for free without ongoing API token costs.
- Training on a local computer without a dedicated GPU takes between 3 to 6 hours per retraining cycle, according to the presenter [07:38].
- Benchmark figures shown: the initial travel model scored 60.28% accuracy, the V2 fine-tuned model achieved 95.28%, compared to Jev's 98.61% on 360 test scenarios (1,440 text decisions) [12:08].
- Another test graphic displays a baseline accuracy improvement from 74.75% before travel training to 93.63% after training across 500 decisions [04:49].
Notable quotes
- "So I took the idea behind Jev and made a version that runs entirely on my computer, completely for free." [00:00]
- "Jev is what's called pretty much a classifier model, specifically it's called an encoder-only model." [02:58]
- "So I wasn't able to quite beat Jev, but I got close enough on a local model running on this computer..." [12:33]
Assessment
This is a technical tutorial and hands-on workflow demonstration. While the web interface, architecture concepts, and prompt engineering methods are shown clearly, long training runs and complete model code execution are abbreviated for presentation purposes.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.