Vritax research log 002 — flight selection

Directional simulation, not a prediction. 120 distinct questions; five repetitions; three description versions; 1,800 measurements.
The earlier multi-intent pilot and all preparation calls are excluded.

questions.csv: exact question strings and generation source.
trials.csv: all 1,800 selection observations, including no-call and diagnostics.
visible-answers.jsonl: visible model output and proposed calls; no tools executed. Outputs may contain errors or unsupported claims; they are observations, not verified travel advice or examples of correct calls.
matrices.csv: all 90 source/application/version aggregates, denominator 100.
paired-deltas.csv: all 90 historical paired comparisons, in percentage points.
original-tool-definitions.json: full original candidate function surface, preserved as presented. All 43 tools were available in the experiment, including non-search tools. Definitions and schema examples are source material, not verified capabilities or real customer records.
descriptions.json: three exact description versions for each changed tool.
study.json: settings, limits and chart data.

Scope: synthetic research questions, model outputs and tool definitions only; no customer conversations or live booking results. Counts describe selection, not successful or appropriate execution.
A predeclared reviewer-coded scope rubric and limited argument checks existed. They were not publisher-validated or a comprehensive expected-action assessment. The published matrices report raw selections; diagnostic flags do not establish correctness.

Reconstruct requests with the question string as input; use tool_choice=auto, parallel_tool_calls=true, truncation=disabled, store=true and explicit prompt_cache_options without breakpoints.
Use the model, reasoning, tier and output cap in study.json. Group original tool definitions by application in their saved order, then concatenate app blocks in trials.csv order.
For detailed/concise, replace only the three matching description strings. No generation-source labels or extra system instructions are sent.
Function names in definitions are the sent aliases; original_name maps them to publisher names. No provider response IDs, credentials, internal filesystem paths or encrypted reasoning are included.
Bootstrap at the question level, preserving repetitions. Observations were made at different times; results are not evidence of a causal length effect.
