Skip to content

Validation and evidence

Claims advance one level at a time. Delivered capability is not observed behavior. Observed behavior is not a measured business outcome. Private evidence is not public proof.

Use this page as the canonical source for evidence status. Service pages may summarize the conclusion, but they must link here and must not strengthen the wording.

MarkerStatusWhat it meansAllowed languageMinimum record
0HypothesisA buyer value, workflow behavior, or outcome is plausible but not yet demonstrated“We expect,” “designed to,” “can influence,” or “target”Named claim, rationale, owner, and next test
1Delivered capabilityThe control exists in a delivered implementation with a defined boundary“FZF built” or “the implementation supports”Scope, acceptance evidence, known limitations, and source
2Observed workflow behaviorThe capability behaved as stated in a live or representative run“In this run…” followed by the exact observationDated run, population, exceptions, logs or records, and recovery details
3Measured outcomeA stable definition and baseline show a result over an agreed period“The measured result was…” with timeframe and attribution limitsBaseline, after-state, method, outside factors, and client acceptance
PPublic proofLevel 1, 2, or 3 evidence may be disclosed in approved wordingOnly the approved claim, context, and underlying levelWritten permission, approved name or anonymization, source link, and review date

Levels 0–3 describe evidence strength. P is a publication overlay, not a stronger causal level. A public capability example does not become a measured outcome merely because it is published.

ClaimCapability deliveredBehavior observedOutcome measuredPublic proofApproved internal interpretation
One run coordinated 80 invoice records; 76 continued while four exceptions stoppedYesYesNoYesSupports exception isolation and controlled stopping in that run
Delivery resumed after an outbound-provider interruption without duplicate sendsYesYesNoYesSupports resumable, duplicate-safe recovery in that run
Staffing Billing Control reduces approved-time-to-delivery or held invoice valuePartlyNot establishedNoNoHypothesis. Establish a buyer-approved baseline and observation period
Staffing Billing Control improves cash, DSO, factoring acceptance, or billing headcountNot established as an outcomeNoNoNoUnsupported. Do not state as achieved or guaranteed

The public source for the observed run is FZF’s staffing invoice automation field note. The four stopped records were three OCR identity mismatches and one known incorrect invoice. The run does not prove a dollar benefit, lower DSO, fewer rejections, faster payment, or headcount savings.

ClaimCapability deliveredBehavior observedOutcome measuredPublic proofApproved internal interpretation
FZF built a control surface for claim state, assignments, recurring obligations, alerts, documents, communications, safety records, reporting, and audit historyYesNot established hereNoNoSupports implementation capability only; source is private
The implementation preserves human review for ambiguous intake, alert resolution, and consequential decisionsYesNot established hereNoNoSupports the designed authority boundary, not a broad production outcome
The service prevents avoidable claim cost, improves return-to-work timing, or increases claims handled per FTENot established as an outcomeNoNoNoHypothesis. Pair direct process measures with contextual outcomes
FZF may name the client or publish a private delivery as a case studyNoNot permitted without explicit written approval

The private delivery includes explicit limits. It has no direct third-party-administrator integration. Other external-system connections and production-hardening work remain bounded by the delivered scope. It does not support claims of regulator certification, predictive outcomes, or autonomous compensability, medical, employment, litigation, reserving, or closure decisions.

  1. Write the smallest true claim. Preserve population, period, environment, and exceptions.
  2. Keep capability, behavior, and outcome separate. “The system routes exceptions” does not prove “the client reduced labor.”
  3. Pair direct and contextual measures. Use approved-time-to-delivery beside DSO, or overdue actions beside claim cost. Do not attribute the contextual result without a method.
  4. Keep null and negative results. They change the scorecard and prevent repeated weak experiments.
  5. Do not aggregate unlike runs. Different clients, rule sets, periods, or definitions need explicit normalization.
  6. Do not borrow permission. Private delivery, repository access, or an internal client name does not authorize public identification.
  7. Time-limit proof. Record the review date and recheck claims when workflow, implementation, or client context changes.

Every experiment must have a short written contract before work starts:

FieldRule
ClaimTest one precise buyer, behavior, or outcome claim
PopulationDefine accounts, records, workflow path, and exclusions
BaselineName the definition, source, owner, period, and known quality limits
InterventionState what FZF changes and what remains outside its control
Primary measureChoose one direct measure that can decide the test
Context measuresTrack broader business results and outside factors without assuming attribution
ThresholdSet the promote, revise, or stop result before observing data
DurationUse a period long enough to include normal exceptions and recovery cases
EvidencePreserve logs, samples, acceptance records, and decision notes
PermissionRecord whether use is private, anonymized, or public

Prefer paid behavior over compliments, observed workflow behavior over stated intent, and client-owned baselines over FZF estimates. In message tests, change one value angle at a time and judge qualified meetings and paid Sprints—not opens, clicks, or generic replies.

Evidence is one part of the portfolio decision. The score and promote/hold/retire rules remain canonical in Portfolio prioritization.

  • The primary buyer, high-consequence job, workflow boundary, and retained authority are explicit.
  • Independent buyer evidence supports the consequence and a usable baseline.
  • At least one core capability reaches Level 1: delivered capability through a prior delivery, prototype accepted against representative data, or paid Sprint output.
  • Integration and recovery risks are known well enough to scope a fixed-price Build.
  • Unsupported outcome language is recorded before sales messaging is approved.

Validated service concept → active sales motion

Section titled “Validated service concept → active sales motion”
  • Paid buying behavior supports demand; interviews alone are insufficient.
  • At least one central workflow claim reaches Level 2: observed workflow behavior or the proposal clearly labels it as a delivery target.
  • The buyer accepts a measurement contract for the first Build or Run period.
  • Delivery capacity, security, authority, and commercial boundaries pass review.
  • Positioning passes the 4 U’s check.
  • The claim reaches Level 3: measured outcome with stable definitions and attribution limits.
  • The client approves the exact wording, identity or anonymization, metrics, period, and channels.
  • FZF stores the supporting source and review date.
  • Public copy states outside factors and avoids guarantees.
  1. Billing: obtain a buyer-approved baseline for approved-time-to-delivery, held value, exception age, completeness, and manual touches; observe the same definitions after implementation.
  2. Workers’ compensation: validate budget and urgency with independent risk leaders; select one bounded path such as injury intake, work-status follow-up, or packet completeness; measure direct behavior before making cost claims.
  3. Proof permission: ask early, but separate permission from delivery acceptance and outcome measurement.
  4. Candidates: spend learning effort in the order and against the unknowns listed in Portfolio prioritization.