Search by meaning, built the way we build for clients.

The right model for the job, our own words on the page, guardrails designed in, and test results published before it shipped.

What happens when you search

  1. 01

    You describe a problem

    In your own words. No keywords to guess, no form to fill in.

  2. 02

    A decision model judges each section

    Jev, a decision model from TypeSafe AI, answers one yes-or-no question for each of the 17 sections of this site and 110 blog posts: is this about what you described? It returns a probability, not prose. One more question picks the best first conversation, or none.

  3. 03

    You see our own words, ranked

    Sections and posts scoring 50% or higher are shown word for word, with links. The ranked results are not written by AI.

  4. 04

    Claude writes the value analysis

    Claude reads only the results Jev chose and explains why the problem matters, where the value would come from, and the evidence from our work, citing each source. It is labeled as written by Claude.

  5. 05

    It fails safely

    If Jev is unavailable, Claude ranks the results instead. If the analysis cannot be grounded in the sources, it is not shown. If nothing answers, the search says so instead of guessing.

Jev decides, Claude writes

Matching a problem to the right part of this site is a judgment, not a writing task. A decision model answers that kind of question in a fraction of a second, returns probabilities that code can act on, and cannot invent claims about our work.

Explaining what a result could be worth to you is a writing and reasoning task, so Claude does that part, from Jev's results only. Choosing the right model for each step, and fencing it in with checks, is part of the engineering.

Tested before it shipped

We ran up to 47 problem descriptions through both engines; the table shows when each engine was last measured and on how many. Most are written to avoid this site's own wording; 6 are off-topic, including one attempt to manipulate the ranking.

MeasureJev (TypeSafe AI)jev-latest · 47 cases · 2026-10-05Claude (Anthropic)claude-opus-5-5 · 45 cases · 2026-10-04
Best first conversation chosen correctly97.3%100%
Top result is a right answer96.8%100%
A right answer in the top three96.8%100%
Right post among the blog results94.1%100%
Off-topic searches correctly return nothing100%100%
Typical time to answer (median)0.28 s13.8 s
Slowest 5% of searches0.38 s16.2 s

Both engines are accurate on this set; Jev answers about fifty times faster, so it runs the search and Claude is the fallback. The set is small and written by us, so treat these as a check, not a benchmark. We rerun it whenever the site's content changes.

Guardrails

  • What you type is not stored and never sent to analytics. We count searches, not words.
  • Your text goes to TypeSafe AI to be scored and to Anthropic to write the analysis (and to rank, if TypeSafe does not answer). Neither trains its models on these API requests under its terms.
  • Your text is treated as data. Instructions inside it are ignored, and the test set includes an attempt to manipulate the ranking.
  • The value analysis may use only the passages Jev returned. Every number in it must appear in those passages or in your own words, or the analysis is withheld. It sizes value by naming what we would measure against a baseline, never by promising an amount, and it says which evidence is a delivered system and which is a prototype.
  • API keys stay on the server. Searches are rate-limited per visitor.
  • Contact-form messages are labeled by the same model so Jerry sees the kind of request first. A person reads every message, and if labeling is slow the message is delivered without a label.

The same pattern, in your workflows

Many steps inside AI workflows are decisions: route this, flag that, does this document answer the question. We find those steps, give each the right model, and measure the result against a labeled set before it goes live.

See how we work

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call