BUSFACTOR.TECH
Buyer’s Guide

We Audited Every 'Deterministic' Engineering Analytics Tool

WE AUDITED EVERY

What we found in vendor docs: LLM classifiers inside scoring loops, undisclosed coefficients, per-developer scores. Every claim cited to the vendors' pages.

15 receipts in this article ↓

AI, audited
Per tool: what shipped, and what got rewritten.

TL;DR: We read the public methodology pages, help docs, deployment docs, and pricing pages of the engineering-analytics field: every major vendor whose marketing leans on "deterministic," "exact," "not survey estimates," or "intelligence you can trust." Three patterns kept repeating: a model (often a literal LLM) inside a load-bearing number, coefficients nobody outside the vendor can audit, and a per-developer composite score. Every claim below links to the vendor's own page. Where we infer rather than quote, we say so. That's the whole point.

The word "deterministic" is having a moment in engineering analytics. One vendor markets "deterministic, not probabilistic" scoring. Another promises attribution built on "usage APIs … not survey estimates." A third brands its research "not self-reported." The category has collectively noticed that engineering leaders are tired of vibes, and it has adopted the vocabulary of auditability whether or not the product underneath earns it.

So we audited the claims. No survey, no hot take: we read what each vendor's own documentation says about how its numbers are computed. This is the write-up, with receipts. Fair warning, it's written by a vendor in this category, which is exactly why every claim here is cited to a public page you can check without trusting us.

What deterministic engineering metrics actually require

The bar is simple to state: same input rows in, same numbers out, every time, plus a published method so someone outside the vendor can re-derive the result. A metric passes or fails on three checks:

  1. No model in the loop. If an ML classifier or an LLM decides what counts - which PR is "AI-assisted," which bucket an engineer's week lands in - the number inherits the model's drift, versioning, and error rate.
  2. Published coefficients. If the formula's weights are proprietary, the number can't be audited, only believed.
  3. Re-runnable history. Re-run last quarter; the output should match byte-for-byte. If the vendor can't offer that test, "deterministic" is a mood, not a property.

Hold every tool to those three, ours included. Here's what the documentation shows.

Pattern one: there's a model inside the number

Navigara is the sharpest case, because its own pages disagree with each other. The ETV methodology post says a language model is used "only to classify ambiguous work into categories, not to invent the number," and that the same diff produces the same score on every run. The AI ROI product page says ETV is "scored straight from commit history by LLMs, ML, and algorithms." And the deployment docs settle the operational question: an LLM backend is required to run the product. Vertex AI by default, AWS Bedrock or Azure AI Foundry as alternatives, with Anthropic Claude models recommended. Note that classification is not a side feature here: the Growth/Maintenance/Fixes split, the bug-fix multiplier, and the CapEx work-type tagging all depend on categories that classification layer produces. To be precise about what's inference on our part: we found no model-pinning, classification-caching, or reproducibility documentation in their public docs as of July 2026. So whether two runs under different model versions produce identical totals is a question their docs don't answer. Ask it.

LinearB launched AI Analytics in February 2026, classifying work as AI-assisted via a confidence threshold. Per their own release notes, the default threshold started at 50, became org-configurable in March 2026, and was cut to 25 in June 2026 "for improved detection accuracy." Read that as a buyer: whether a PR counts as AI-assisted is a tunable probability, and the vendor moved the default cutoff mid-year, which quietly changes what every earlier adoption chart meant. The scoring model that produces the confidence value is not disclosed.

Jellyfish builds its flagship on the Work Model, which "ingests, cleans, normalizes, and scores signals" from Jira and Git to split each engineer's week into FTE fractions across projects. That is the number finance then prices with payroll. The scoring function is patented and not disclosed, and the current product pitch adds AI-powered categorization on top. No reproducibility guarantee appears anywhere in their public material.

Swarmia deserves the most nuance. Its AI-detection docs disclose four signals with stated confidence tiers: three deterministic (commit authorship, git trailers, PR labels) and a fourth, author-used-an-AI-tool-within-24-hours heuristic that their own docs grade low-confidence and admit "can result in overreporting." That heuristic is on by default in AI-assisted counts, and a separate AI auto-categorizer sits inside Investment Balance - the number that feeds capitalization reports. Honest disclosure, genuinely. But the probabilistic layer is inside the load-bearing figure.

Faros AI brands its research "not self-reported", then runs Lighthouse AI, "statistical analysis, machine learning, and GenAI," as the insight layer that summarizes dashboards and answers questions. The rows may be telemetry. The interpretation an executive reads is generated.

Pattern two: coefficients you can't audit

Even where no LLM is involved, several flagship numbers rest on arithmetic nobody outside the vendor can check.

  • DX computes the DXI as a composite of 14 Likert survey items and translates it to money with a striking claim: one DXI point equals 13 minutes saved per developer per week. The regression behind that coefficient is not published. You can cite the number, but you cannot re-derive it.
  • GitClear publishes more of its Diff Delta mechanics than anyone (base scores per line operation, a filter cascade, context weighting) and still keeps the final scalars and churn windows proprietary. The best disclosure in the field is still not a formula you can run.
  • LinearB's benchmark bands come from a genuinely impressive corpus - 8.1M+ PRs across 4,800 teams, per their published benchmarks report
    • but the percentile math behind the band cut-points isn't disclosed on the report page.
  • Navigara names five weighting factors for ETV and publishes zero weights.
The AI view comparing coding tools by volume shipped and code later rewritten, with an honest below-sample row for the tool with too few pull requests to judge.The AI view comparing coding tools by volume shipped and code later rewritten, with an honest below-sample row for the tool with too few pull requests to judge.
The AI-tool compare - shipped versus rewritten, per toolLive product · fictional demo org

Pattern three: the per-developer score

The third finding is less about determinism than about what the deterministic machinery gets pointed at. A remarkable share of the field ships a per-person composite score:

  • GitClear computes Diff Delta per developer: a productivity credit score per human.
  • Harness documents Trellis Scores, a per-developer composite with org-adjustable factor weights, and its AI DLC launch blog describes an on-machine agent capturing full session data, including individual conversations and prompts.
  • Navigara benchmarks ETV "across teams, repos, and developers" and publishes org-average ETV-per-developer figures for outside organizations on its landing page.
  • Typo surfaces "AI champions" and per-individual AI usage and cost.
  • Even Multitudes, whose ethics we genuinely respect, auto-labels each individual a High or Low AI adopter using a 35-percent-of-days threshold over a 12-week window. Person-grain classification, transparently disclosed, but person-grain all the same.

A sortable per-developer score column is a layoff spreadsheet with extra steps, whatever the vendor's intent. The research case against the whole pattern is laid out in should you measure individual developer performance and why stack ranking backfires. The short version: a single number per human measures neither productivity nor value, and your senior engineers will game or resent it within a quarter.

Credit where it's due

An honest audit concedes what it finds. Swarmia's confidence-tier disclosure is the most honest AI-measurement documentation among the majors. Multitudes publishes caveats most vendors would bury, like selection bias in adopter cohorts and inconsistent impact across metrics. GitClear publishes more methodology than any direct rival. And DX's own research states that realistic AI productivity gains run five to fifteen percent rather than the fifty-to-hundred the hype cycle sells - a finding worth quoting against AI-tool vendors, from a measurement vendor. None of that rescues an undisclosed coefficient or a default-on heuristic, but it beats the alternative: several vendors in this field disclose essentially nothing about attribution mechanics at all.

The delivery-stats view breaking each pull request into pickup, review, merge, and deploy time.The delivery-stats view breaking each pull request into pickup, review, merge, and deploy time.
The cycle-time breakdown - where each PR spends its lifeLive product · fictional demo org

The test to run on any vendor, including us

If you're evaluating this category, the full question list covers the ground; the determinism-specific cut is five questions:

  1. Re-run last quarter. Does the output match byte-for-byte? Will you demo that?
  2. Is any model, ML or LLM, anywhere in the metric path, as opposed to the chat UI?
  3. Are the formula's coefficients published, fully, somewhere I can link?
  4. What changed about historical numbers the last time you retuned a default?
  5. Does any surface compute a composite score per person?

Busfactor's answers, on the record: our metric path runs zero LLMs. Numbers are computed from your rows and quoted, never generated. Re-runs are byte-identical, and exports print the provenance (run id, engine version, ruleset version, content hash) so the re-run test is checkable rather than promised - the same posture behind the deterministic verdict layer and the finance-grade money views. People get retention-framed scorecards, and the one ordered view - Standing - prints its arithmetic and its receipts on every position. What we don't have, equally on the record: our benchmark bands are published, cited reference points rather than a live peer cohort (some rivals genuinely beat us there), and our AI attribution deliberately counts only what can be quoted (trailers, bot accounts), with the blind spots disclosed, as explained in how to measure AI adoption.

The bottom line

The market has learned that CTOs want auditable numbers, and the marketing has adapted faster than the architectures. When a vendor says "deterministic," the truth lives in the docs rather than the homepage: look for the required LLM backend, the tunable threshold, the patented scoring function, the unpublished regression. Estimates disclosed as estimates are fine. Estimates sold as facts are how you end up defending a number in front of your CFO that nobody on earth can recompute. The honest buyer's guide covers the rest of the purchase. On determinism specifically, trust the tool that invites the byte-diff test, and check whether AI is making your code worse with numbers you can actually re-run.

Frequently asked

What does deterministic mean for engineering metrics?

Same input rows in, same numbers out - every time, to the byte. That requires a published method, no probabilistic model anywhere in the computation, and a way to re-run a past period and get identical output. If any step involves an LLM classification, a tunable confidence threshold, or coefficients the vendor won't publish, the number is an estimate wearing a determinism costume - which can still be useful, as long as it's labeled.

Do engineering analytics tools use AI to compute metrics?

Increasingly, yes - and not just in chat assistants. Several vendors run classifiers inside load-bearing numbers: which work counts as AI-assisted, which bucket a week of effort lands in, which costs get capitalized. The documentation is the tell: look for words like 'confidence threshold,' 'AI-powered categorization,' or a required LLM backend in the deployment docs.

Is a probabilistic metric always bad?

No - estimates are legitimate when disclosed as estimates. Swarmia labels its low-confidence AI-detection signal as such, and Multitudes publishes its caveats, and both deserve credit for it. The failure mode is marketing a probabilistic pipeline as 'deterministic' or 'exact' while the formula, the model version, or the threshold stays undisclosed and movable.

Receipts

Keep reading