A GPU experiment checklist for AI research

A training job can finish without answering the research question. Use this checklist before allocating compute and again when reviewing returned results. It is a review aid for researchers, not a claim that any particular experiment has passed.

Before starting a run

  1. Write the hypothesis, prediction target and conditions under which a result would support or contradict it.
  2. Record the dataset version, source and permissions; identify which inputs exist at prediction time.
  3. Document training, validation and test splits. Consider separation by device, subject or time where the task requires it.
  4. Fit learned preprocessing on training data and apply it consistently to the other splits.
  5. Choose the primary metric, its units, aggregation rule and whether a larger or smaller value is better.
  6. Fix comparable baselines, tuning rules and training budgets before inspecting final test results.
  7. Define ablations, random seeds, the number of runs and a stopping rule within the available compute budget.
  8. Check that the planned run will return the configuration, execution record and outputs needed to verify the reported metrics.

Separate evaluation from model selection

Use validation results to select model settings. Keep final test results out of that selection process. For repeated measurements from a machine or person, inspect whether near-duplicate observations cross splits. For forecasting, check whether feature construction accidentally uses future information. A split should reflect the setting the research claim is meant to describe.

Make comparisons answer a specific question

A baseline comparison asks whether the proposed method improves on a relevant alternative under a documented protocol. An ablation asks what changes when a particular component is removed or replaced. Keep the other settings comparable and record unavoidable differences. More parameters, extra training data or a larger tuning budget can change the interpretation of an apparent gain.

Record variation as well as the best run

When multiple runs are planned, preserve each result and explain how the summary is calculated. If only one run is affordable, say so instead of implying repeated confirmation. Record software versions and hardware alongside seeds: a seed alone does not guarantee identical results across different environments.

Evidence to inspect after the run

Use this table as a handoff checklist. For each row, record the actual file or record location and any unresolved gap in the project notes. These are requested evidence types, not guaranteed files or sample results.

Evidence to inspect after the run
CheckEvidence to locateWhat it should establish
Data and splitDataset version, split rules and sample or group countsWhich observations were used, and whether the evaluation matches the intended setting.
Method and baselineExecuted code version, model configuration and tuning protocolWhat was compared and which settings differed.
ExecutionRun status, logs, environment and saved configurationWhether training and evaluation completed; completion alone does not establish scientific validity.
MetricsReturned metric summary and the corresponding evaluation outputsMetric definition, units, split, aggregation and linkage to the evaluated run.
Ablations and repeatsPer-condition and per-run results, including failed runsWhether the claimed effect is supported across the comparisons actually completed.
FiguresSource values and plotting configurationWhether plotted values, axes, labels and captions match the recorded evidence.

When a result is not ready to support a claim

An absent output, ambiguous split, unmatched configuration or unfinished comparison is a gap to resolve. Label the result as incomplete and limit the conclusion to what the available evidence supports. Do not replace missing runs with estimated scores or select only favorable outcomes.

Using the checklist with AutoResearch

AutoResearch connects model implementation with GPU training, evaluation, ablations and inspectable returned artifacts. The checklist helps researchers review those artifacts and decide what still needs to be verified. It does not certify a result or replace domain expertise.

Technical references