Updated September 6, 2026 · Evaluation protocol, not a results report

Judge the work at the level the evidence supports.

Current public evidence

We do not publish an authorized customer-generated artifact, human comparative study or measured customer uplift dataset on this site. The public walkthrough is manually authored and explicitly illustrative. It cannot establish output quality, launch reliability, time savings or revenue impact.

1. Freeze a real, permitted brief.

Before generation, record the version and date of the business brief, target audience, offer, goals, approved claims, forbidden claims, channel, budget constraints and source material. Use only data you have permission to process. Keep private customer data out of public evaluation materials.

Record missing information instead of asking a model to fill it with plausible facts. Define the acceptance criteria before viewing output, so the criteria cannot be changed to favor a result.

2. Preserve the run and the revision.

For an actual product evaluation, retain the project/job/artifact identifiers, execution mode, provider/model identifiers where available, generation timestamps, failed attempts, retries and revision history. Inspect the final artifact alongside the exact brief and source versions used.

Separate initial output from post-edit output. Include failures and blocked runs in the denominator. A redacted assurance export can support inspection only to the extent that its included evidence permits; missing logs or unavailable configuration must be disclosed.

3. Review quality without turning a model score into truth.

Genesis can attach quality findings and perform deterministic and model-based checks. Available critics and review paths depend on configuration. Multi-provider agreement does not establish factual accuracy, independence in the scientific sense or campaign effectiveness.

Use qualified human reviewers for factual, brand, legal and domain judgment. Record reviewer identities or roles, conflicts of interest, rubric version and disagreements. AI scores and AI critics' verdicts must be labeled as such—not attributed to human agencies or independent customers.

4. Evaluate a comparator fairly.

If comparing with an existing team process or another tool, use the same permitted brief, source material, delivery requirements and time accounting. Randomize or blind presentation where practical. Specify the unit of evaluation, sample selection, number of briefs/reviewers, scoring criteria and treatment of failed or unfinished runs in advance.

Report distributions and uncertainty, not only the best example. Include subscription, metered/provider cost, retries and human correction time. Publish the conditions and limitations with any result. The website comparison summarizes dated official product descriptions; it is not such an experiment.

5. Prove execution separately from preparation.

  1. Draft: inspect the actual stored content and parameters.
  2. Approved: inspect authorization for that revision. Verify current consent, account ownership, budget and safety controls.
  3. Queued: identify the accepted job and its status. A local offline queue is not server acceptance.
  4. Provider accepted: retain an authentic provider identifier or response with a timestamp. This does not prove delivery.
  5. Delivered or published: reconcile the appropriate delivery callback, provider status or accessible destination. A positive API response alone is insufficient.
  6. Measured: link eligible outcome records to the relevant campaign and time window. Distinguish deduplicated conversions from repeated events and attributed revenue from incremental profit.

A signed export establishes integrity of the exported bytes under a trusted verification key. It does not prove source truth, universal capture, lawful processing or a correct decision.

6. Measure outcomes before claiming uplift.

Define the eligible population, assignment unit, primary outcome, measurement window and analysis plan in advance. A causal claim needs an appropriate identification design, such as a randomized holdout that is actually excluded at the send boundary—not a simulated control group or correlation score.

Check assignment balance, exposure, contamination, duplicate events, attribution windows and missing data. Report uncertainty and sample sizes. Use consistent monetary units and include relevant spend and costs. If outcome tracking, comparable controls or sufficient data are absent, report “unavailable” or “inconclusive,” not zero or positive lift.

Search visibility requires identifiable real probes, retrieved responses and explicit brand/domain matching. Simulated engine answers are research previews, not observed ranking or share of voice. Model confidence is not causal confidence.

7. Publish only with permission and attribution.

A publishable case needs written permission for its artifacts and customer identity, appropriately redacted sources, the evaluation protocol, dates, complete denominators and limitations. No private brief, contact list, credential, raw customer export or provider receipt should become a public fixture by default.

Until that evidence exists, this site makes no agency-superiority, fixed productivity-multiple, guaranteed launch-time or revenue-uplift claim.

Start an evaluation in your workspace.

Create a real project brief or review existing decisions. These routes require sign-in, an eligible plan and workspace authorization. New to UltimateNexus? Request early access to discuss fit and prerequisites.