A successful AI demo shows a selected interaction. A reproduction report asks whether the same claimed task can be performed again under recorded conditions, and explains where the rerun differs from the original presentation. That report can be useful even when the outcome is mixed. Its value comes from the clarity of the task, the completeness of the record, and the restraint of the conclusion.
This guide provides a practical reporting protocol for AI editors, with primary technical documentation checked on October 6, 2026. It does not report a live reproduction of a particular commercial demo. The sample tasks and run outcomes are explicitly hypothetical examples of how to structure a test; they are not performance findings about Gemini, PyTorch, or another provider.
Define success before pressing run
Extract one observable claim from the demonstration. If a video shows a model reading a photograph, naming an object, and creating a shopping list, choose whether the rerun targets object identification, list generation, or the combined workflow. Write the expected input and outcome in a sentence that another reviewer could apply without guessing what the presenter meant.
Decide whether you need an exact output match or a functionally equivalent result. For a task that names three objects, different sentence wording can still meet the criterion. For a task that returns a required data structure, the exact fields may matter. For a task that claims a particular response time, define the measured interval and include the time needed to complete the relevant action.
Keep acceptance rules independent of the result. If an answer names two objects correctly and misses the third, do not redefine success afterward to mean identifying βmost objects.β Record the partial result and explain it against the original criterion. This protects the report from becoming an account of whichever run happened to look most persuasive.
Freeze the input and the surrounding context
Retain the original prompt text and the exact input files when they are available and you have permission to use them. Record file names, formats, dimensions, and a content hash for local fixtures. A hash helps identify whether two runs used the same bytes; it does not reveal whether the file represents the complete input used in the original demonstration.
If the original input is unavailable, create a clearly identified substitute and describe the difference. A new photograph of the same type of object is an analogous test, not an exact reproduction. A transcript of spoken audio leaves out vocal information. A screenshot from a video may omit earlier frames that supplied context. These differences should appear in the report's comparison section.
Capture the surrounding state as well. Previous conversation turns, uploaded documents, account memory, connected tools, and system instructions can affect the task context. Start a fresh session when the protocol calls for one, and say whether a session was fresh or reused. Unknown original state should remain an explicit limitation rather than being silently treated as empty state.
Match the product route and input handling
Use the same app or API route as the demonstration when possible. A consumer interface can supply tools or processing that a plain model request does not include. A hosted service and a local checkpoint can also differ in ways that matter to the task. Record the route, exact model identifier if exposed, product version, account conditions, and date of the run.
For an image task, consult the relevant primary input documentation before treating a failed upload as a capability failure. Google's Gemini API image-understanding guide documents supported image inputs and their handling. It lists PNG, JPEG, WEBP, HEIC, and HEIF formats and provides image-input examples. This is evidence about that documented API route, not proof that a particular image task will succeed.
A reproduction record should say whether the image was supplied as original bytes, through a file reference, or through another documented mechanism. Also note any resizing, cropping, rotation, or conversion performed before submission. Such preparation may be necessary, but it changes the input and should be visible to anyone repeating the test.
Control what the environment lets you control
Record exposed generation settings, output limits, tool choices, and any seed value. If the interface does not reveal a setting, mark it as unavailable rather than inventing a default. For local tests, retain the code revision, model checkpoint, framework version, dependencies, hardware, and relevant execution configuration. For hosted tests, retain the observable request and response metadata without implying access to hidden infrastructure.
PyTorch's official reproducibility notes warn that complete reproducibility is not guaranteed across releases, commits, or platforms, and that CPU and GPU runs can differ even with identical seeds. For a local PyTorch experiment, this supports recording the environment alongside the seed. It should not be presented as a claim about the implementation of an unrelated hosted service.
Separate exact reproducibility from meeting a task criterion. Two outputs can differ while both correctly complete the task. Conversely, identical outputs can repeat the same mistake. If the claim concerns factual accuracy or a completed action, evaluate that outcome directly instead of treating repeated wording as proof of success.
Predefine attempts and maintain a complete log
Specify the number of planned attempts, whether sessions will reset, and how retries will be handled. A small exploratory rerun can be informative, but it is not a broad reliability study. Choose an affordable test scope and state its limits rather than implying that a few favorable attempts establish performance for all users or inputs.
Create a row for every attempt, including failures and incomplete runs. Record the initial output before rewriting the prompt or adding a missing file. If a repair is needed, log it as a separate intervention. A corrected second attempt can show that the workflow eventually succeeded with assistance; it should not replace the unsuccessful first attempt in the record.
| Run-log field | What to retain | Why it matters |
|---|---|---|
| Task and fixture | Criterion, prompt, input reference, and hash | Makes the attempted task identifiable |
| Environment | Route, model label, version, settings, and session state | Shows what conditions were controlled |
| Raw outcome | Original response and resulting artifact | Preserves evidence before interpretation |
| Task assessment | Pass, partial result, failure, or not assessable | Separates the output from the verdict |
| Interventions | Retries, prompt edits, manual help, and tool changes | Reveals how the final result was obtained |
| Operational notes | Errors, timing definition, and observed costs | Limits conclusions about the workflow |
Distinguish a task failure from a blocked test
A task failure occurs when the system receives the intended input and does not meet the defined criterion. A blocked test occurs when you cannot submit the input, obtain the required access, or reach the relevant service. Both belong in the record, but they support different conclusions. An account denied access cannot establish whether an inaccessible capability would have completed the task.
Partial execution also needs careful treatment. An assistant might generate a correct instruction but never execute the promised action. It might create a file that fails to open. Verify the final artifact or action against the criterion instead of relying on the assistant's own statement that the work is done. Where verification requires an external system you cannot inspect, mark the action as unverified.
If an operational error affects timing, explain whether the attempt remains relevant to the claim. A rerun after an outage can illuminate capability under restored conditions, while the outage remains part of the observed user experience. Keep these findings separately described rather than dropping inconvenient operational records or treating every error as model reasoning failure.
Interpret an illustrative run record
Consider a hypothetical demo that claims to identify three visible objects in a single image without follow-up questions. A newsroom plans four fresh-session attempts using the same fixture and criterion. In this illustrative record, two attempts name all three objects, one omits an object, and one never completes because the request is rejected before generation. No real provider results are represented by these invented outcomes.
The report would describe two successful task completions, one partial result, and one blocked attempt. It should retain all four rows. If the writer gives a proportion, its denominator and treatment of the blocked attempt must be explicit. Reporting only the two successful screenshots would conceal the task's observed inconsistency within the example.
Suppose a follow-up prompt corrects the omission. Log that result as an assisted continuation. It may be useful evidence that interaction improves the outcome, but the original one-shot criterion was not met in that attempt. Likewise, replacing the fixture with a cleaner image establishes another test condition rather than repairing the historical record of the first run.
Publish a conclusion that someone else can repeat
Open the reproduction report with the tested claim, exact scope, and observed outcome. Identify the date, route, fixture, acceptance rule, attempt count, and material deviations from the original demo. Include links to the relevant primary documentation near technical statements. When possible, provide the prompt and a safe, reusable fixture so that another reviewer can examine the same task.
Use a limited conclusion such as βthe task criterion was met in these recorded attempts under these conditions.β If no matched run was possible, say that reproduction remains untested or blocked. Avoid extending a small sample into βthe product always works,β or using an unmatched failed task to declare that the original demonstration was impossible.
Ask a reviewer who did not design the test to inspect the acceptance rule and raw record. They should be able to see why each attempt received its assessment and where intervention occurred. When product versions change, add a new dated test series rather than overwriting the earlier findings. The research-to-product checklist helps establish which system the demo represents, while the benchmark-conditions guide covers broader comparisons across evaluation setups.
Frequently asked questions
Must the rerun produce exactly the same words?
Only when exact wording is part of the tested claim or required output. Otherwise, define a functional criterion, such as correctly identifying specified objects or producing a valid file. Different wording can meet the same criterion, while identical wording can reproduce an error.
What if the original prompt or file is unavailable?
Record that limitation and identify any substitute input explicitly. You can test a similar capability under your own documented conditions, but should not call it an exact reproduction. Explain which differences could affect the comparison with the presented demonstration.
Does setting a seed guarantee identical results?
Do not assume that it does. Check the relevant system's documentation and record the environment as well as exposed settings. For local PyTorch work, its official notes identify reproducibility limits across versions and platforms. A hosted interface may expose different controls or no seed at all.
Should upload errors count as failed model answers?
Record them as operational or blocked attempts when generation did not receive the intended input. Keep them in the overall run log, but distinguish them from assessed task outputs. State how any reported count or proportion treats these attempts so the reader can interpret it correctly.
How many attempts prove a demo is reproducible?
There is no universal count that proves reliability for every task. Predefine a scope suitable for the claim and resources, retain all attempts, and limit the conclusion to the tested conditions. Larger claims about reliability require a more extensive evaluation than a small editorial rerun.
