Expected Parrot · A practical, evidence-first tutorial

Creating and validating digital twins from survey data

Build twins from respondents’ known answers, hide a target question, predict its probability distribution, and test whether the result contains individual-level signal beyond cheap baselines.

Tool: zwillRuntime: EDSLExample: Pew Wave 158

01 What is being tested?

A survey digital twin is an LLM stand-in for one respondent. Its prompt contains approved facts about that person—other survey answers and readable covariates—but not the answer it must predict. The model returns probabilities over the held-out question’s options.

\hat p(yiq | xi,-q)Predict respondent i’s answer to question q using their information other than q.

Observed truth

The respondent already answered the target. That answer is hidden from the twin and retained for scoring.

Twin prediction

A full probability vector—not only a top choice—so uncertainty and calibration can be evaluated.

Cheap comparison

A conditional model predicts the same targets from the same allowed respondent information.

Decision gate

The twin must beat the conditional baseline with uncertainty that clears zero, under a clean leakage audit.

The key distinction. Reproducing a population marginal is not the same as predicting individuals. A twin can sound plausible and match aggregate opinion while adding no respondent-level signal.
Import observed survey
Hold out target answers
Run one-shot + twins
Audit prompts + score
Gate the claim
zwill preserves the artifacts at every boundary so the result can be reproduced and audited.

02 Study design before model calls

Our worked example uses a fixed excerpt of 100 respondents from Pew Research Center’s American Trends Panel Wave 158. The survey asked whether people favored or opposed six climate-policy proposals. We will hide each person’s answer about taxing corporations based on their carbon emissions and try to predict it from their other five policy answers.

100 respondents × 1 target × 1 twin approach × 1 model = 100 model calls
ComponentChoice in this tutorialReason
Targetcarbon_taxIts 62–38 unweighted split is less lopsided than several other battery items.
Twin contextThe respondent’s other five climate-policy answersRelated individual evidence is available without exposing the target.
Validation sample100 fixed respondentsEnough to fit and demonstrate XGBoost while keeping remote inference affordable.
Conditional baselineXGBoostTests whether an LLM adds value beyond a conventional learner using comparable evidence.
Primary scoreNLLDeclared before inference because it evaluates the full probability distribution and strongly penalizes confident errors.
Secondary scoresBrier, p(actual), and top-1Useful corroborating views, but they do not replace the declared primary comparison.
This is a real-data pilot, not a representative subsample or definitive validation. The excerpt is the first 100 normalized source records, not a probability sample drawn for this analysis. Retaining Pew’s weights does not repair that convenience selection. A publishable claim should use the documented full sample, rotate across several held-out questions, preregister the design, and account for trying multiple models or prompts.

03 Installation and orientation

This tutorial assumes you have cloned the zwill repository and are standing at its top level—the directory containing README.md, pyproject.toml, and examples/. A deidentified 100-person teaching excerpt is already present under examples/pew_w158_tutorial/, so every data path below can be copied exactly.

git clone https://github.com/expectedparrot/zwill.git
cd zwill
pip install -e ".[conditional-baseline]"

zwill guide
zwill next

The conditional-baseline extra installs XGBoost and a local sentence-transformers embedding option. zwill guide is the source of truth for the complete workflow. zwill next inspects local state and returns the next command after every completed stage.

Where files go: zwill places relative report and export paths beneath its default output root, zwill_work/. Thus --path survey-profile.html writes zwill_work/survey-profile.html, and --out report_out/validation writes zwill_work/report_out/validation/. EDSL interchange files such as jobs.ep and results.ep stay at the paths written in the command so the external ep CLI can find them. Set ZWILL_OUT or use zwill init --output-dir to choose another output root.

Model runs require an EXPECTED_PARROT_API_KEY in a nearby .env. If you do not have one, sign up with Expected Parrot. zwill prepares durable EDSL .ep job packages; EDSL’s ep CLI runs them and writes a Results package. This keeps credentials, remote execution, waiting, and retries in the inference tool rather than the validation tool.

The execution boundary. zwill builds and imports artifacts; it does not run models. Inspect a job with ep inspect jobs.ep, execute it with ep run jobs.ep --output results.ep, and reopen the result in Python with Results.git.load("results.ep").

04 Ingest the Pew survey excerpt

The excerpt contains six questions, 100 deidentified respondents with survey weights, and 600 observed answers. Source response codes are expanded to the canonical labels Favor and Oppose. Coded demographic fields were deliberately omitted because their codebook is not bundled here; neither readers nor models should be asked to guess what opaque codes mean.

Create zwill’s local project database and the empty survey record that later commands will address.

zwill init
zwill survey create --name pew_w158

Retain the source evidence

A trustworthy analysis keeps source and provenance notes beside the normalized records. We retain this material both for auditability and so an agent can reason over the original study context later when normalized rows alone are insufficient. raw add copies the note into managed state and records its hash, title, type, and original path; it does not automatically place the text into every model prompt.

zwill raw add --survey pew_w158 --id pew_w158_excerpt \
  --input-path examples/pew_w158_tutorial/source.md \
  --kind documentation \
  --title "Pew American Trends Panel Wave 158 excerpt"

The excerpt is included for reproducible teaching, criticism, and methodological validation. Before publishing a substantive analysis, consult Pew Research Center’s Wave 158 dataset page, survey methodology, and linked questionnaire and access terms. Pew Research Center bears no responsibility for this tutorial’s analysis or interpretations.

Add one question, respondent, and answer

Before bulk import, it helps to see the three connected records directly. The question establishes wording and legal options; the respondent establishes an ID and survey weight; the answer links that person to one declared option.

Use a disposable survey name for this small manual demonstration so the complete import below remains directly runnable.

zwill survey create --name pew_w158_manual

zwill question add --survey pew_w158_manual \
  --question-name carbon_tax \
  --question-type multiple_choice \
  --question-text "Do you favor or oppose taxing corporations based on the amount of carbon emissions they produce?" \
  --question-option Favor --question-option Oppose \
  --role survey_item --source-raw pew_w158_excerpt \
  --source-note "Wave 158 CCPOLICY battery item b"

zwill respondent add --survey pew_w158_manual \
  --respondent-id w158_ccpolicy_100314 \
  --weight 0.4510313485401432 \
  --source-raw pew_w158_excerpt

zwill answer add --survey pew_w158_manual \
  --respondent-id w158_ccpolicy_100314 \
  --question carbon_tax --answer Favor
What just happened? Favor is valid because it was declared as an option before the answer arrived. The weight says how much this respondent contributes to weighted population summaries. The deidentified ID is only a stable join key; it is not evidence supplied to the twin.

The three kinds of normalized record

Public-use microdata usually arrives as one wide row per respondent. zwill separates that table into three JSON Lines files so its assumptions can be checked explicitly.

RecordWhat it describesKey validation
QuestionStable name, exact wording, type, options, role, and provenanceOptions are declared before answers arrive
RespondentDeidentified ID, survey weight, and provenanceEvery answer refers to a known respondent
AnswerOne respondent’s answer to one questionThe value belongs to that question’s declared option set

JSONL stores one complete JSON object per line without a surrounding array. Representative lines from the bundled files look like this:

{"question_name":"carbon_tax","question_type":"multiple_choice","question_text":"Do you favor or oppose ... taxing corporations based on the amount of carbon emissions they produce","question_options":["Favor","Oppose"],"role":"survey_item","source":{"raw_id":"pew_w158_excerpt","note":"Source codes expanded through the metadata codebook."}}

{"respondent_id":"w158_ccpolicy_100314","weight":0.4510313485401432,"source":{"raw_id":"pew_w158_excerpt","note":"Deidentified source respondent."}}

{"respondent_id":"w158_ccpolicy_100314","question":"carbon_tax","answer":"Favor"}

Import the complete 100-person excerpt

The manual records went into pew_w158_manual; the empty pew_w158 survey created at the start is therefore still ready for the complete import. Import all six questions, 100 respondents, and 600 answers:

zwill question import --survey pew_w158 \
  --input-path examples/pew_w158_tutorial/questions.jsonl
zwill respondent import --survey pew_w158 \
  --input-path examples/pew_w158_tutorial/respondents.jsonl
zwill answer import --survey pew_w158 \
  --input-path examples/pew_w158_tutorial/answers.jsonl

zwill next --survey pew_w158

Import is a validation boundary, not a best-effort conversion. Unknown respondents, undeclared questions, and illegal options are reported rather than silently changing the survey. Because the excerpt is complete by construction, every respondent should have six accepted answers and there should be no open quarantine issues.

05 Inspect, then commit observed truth

Run each command separately and read what it returns. Start with project status to confirm that the expected records arrived and that the survey has not yet been committed.

zwill status

Output:

{
  "command": "zwill status",
  "status": "ok",
  "data": {
    "project": "default",
    "surveys": [{
      "name": "pew_w158",
      "status": "draft",
      "raw_files": 1,
      "questions": 6,
      "respondents": 100,
      "answers": 600,
      "open_quarantine_issues": 0,
      "committed": false
    }]
  },
  "next_steps": ["zwill commit --survey pew_w158"]
}

Next inspect the respondent-by-question table. This is the fastest way to catch shifted columns, untranslated codes, or unexpected missing values.

zwill table --survey pew_w158 --limit 5

Output shape:

                              pew_w158 answers
┏━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ respondent_id ┃ tree_planting ┃ carbon_tax ┃ carbon_captu… ┃ zero_carbon_… ┃ methane_leaks ┃ home_effici… ┃
┡━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ w158_…_100314 │ Favor         │ Favor      │ Favor         │ Favor         │ Favor         │ Favor         │
│ w158_…_100363 │ Favor         │ Favor      │ Favor         │ Oppose        │ Favor         │ Favor         │
│ w158_…_100637 │ Favor         │ Oppose     │ Oppose        │ Oppose        │ Favor         │ Oppose        │
└───────────────┴───────────────┴────────────┴───────────────┴───────────────┴───────────────┴───────────────┘

Create a fuller HTML profile to inspect exact wording, weights, missingness, and weighted response distributions.

zwill survey report --survey pew_w158 --format html \
  --path survey-profile.html

The response reports the written path and the six-question, 100-respondent, 600-answer dimensions. Open zwill_work/survey-profile.html before committing and verify that every item is binary Favor/Oppose and has 100 observed responses.

commit freezes the respondents and truth marginals used later for scoring. It is not merely a save button.

zwill commit --survey pew_w158

Output:

{
  "command": "zwill commit",
  "status": "ok",
  "data": {
    "survey": "pew_w158",
    "respondent_count": 100,
    "question_count": 6,
    "answer_count": 600,
    "truth_marginal_count": 6
  }
}

Now ask zwill what stage follows from the committed state:

zwill next --survey pew_w158

It reports run_twins because the observed truth is committed but no held-out twin run has been imported yet.

FieldRole in validationTreatment
carbon_taxHeld-out targetWording and options shown; respondent’s answer hidden
The other five policy itemsRespondent contextAnswers shown to both the twin and conditional baseline
Survey weightPopulation estimationUsed for weighted summaries, never shown as personality evidence
Respondent IDStable join keyUsed to pair predictions with truth, not as predictive evidence

06 Establish the one-shot marginal

Before constructing person-specific twins, ask a model for the population distribution of carbon_tax. This one-shot task contains the question and its two legal options but no respondent. It tests whether the model can approximate aggregate opinion from general knowledge alone.

1. Build the probability job

zwill edsl build --survey pew_w158 --target probability-job \
  --questions carbon_tax --model openai:gpt-5.5 \
  --path one_shot_jobs.ep

This creates the local execution specification. It does not contact the model.

2. Inspect the job

ep inspect one_shot_jobs.ep

Output:

{
  "status": "ok",
  "data": {
    "object_type": "Jobs",
    "length": 1,
    "question_count": 1,
    "question_names": ["response_probabilities"],
    "agent_count": 0,
    "scenario_count": 1,
    "model_count": 1
  },
  "warnings": []
}

There is one scenario because this is one population-level question, not 100 respondent agents. Inspection lets you catch the wrong target, model, or job size before paying for inference.

3. Run it with EDSL

ep run one_shot_jobs.ep --output one_shot_results.ep

The ep CLI performs remote inference and stores the completed EDSL Results object at the requested path.

4. Import the result

zwill prob-results import --survey pew_w158 \
  --input-path one_shot_results.ep

The tutorial run imported one of one rows with no issues and cost approximately $0.03. Its response returned data.job_id: a4399bc71f1b07e7. Your run receives its own ID; copy it wherever the tutorial shows <probability_job_id>.

5. Plot predicted versus observed marginals

A marginal is the overall share choosing each option, without retaining which respondent chose it. In this excerpt, 62 of 100 respondents favored the carbon tax and 38 opposed it. After applying the bundled survey weights, the observed shares are 60.8% Favor and 39.2% Oppose.

Zwill aligns the two probabilities returned by the language model with the ingested canonical options Favor and Oppose, then plots them beside those weighted observed shares:

zwill prob-results report --survey pew_w158 \
  --job-id <probability_job_id> --format svg \
  --path one-shot-marginals.svg
Observed and GPT-5.5 one-shot Favor and Oppose shares for the Pew W158 carbon-tax question
The actual tutorial run predicted 69% Favor and 31% Oppose, compared with the weighted observed split of 60.8% and 39.2%. The model was closer than a 50–50 uniform prediction, but overstated support by 8.2 percentage points.

Open your generated zwill_work/one-shot-marginals.svg. The observed bars come from committed microdata; predicted bars come from the imported model result. This run showed useful population-level knowledge, but it does not establish that the model can distinguish one respondent from another.

What this baseline cannot establish. A model can predict 61% Favor for everyone and reproduce the sample marginal while offering no individual signal. The held-out test in the next chapter asks the harder question.

07 Run held-out predictions

We now move from “Can a model guess the aggregate split?” to “Can it assign better probabilities to each real person?” The committed survey already contains every respondent’s carbon_tax answer. Zwill keeps that answer locally for scoring but removes it from the scenario sent to the model.

Model may seeModel must not see
The exact carbon_tax wording and its Favor/Oppose optionsThe respondent’s recorded carbon_tax answer
The respondent’s answers to the other five climate-policy proposalsTarget summaries or derived fields that reveal the hidden answer
General survey context explicitly approved for the studySurvey weight or respondent ID presented as personality evidence

Write and review the validation plan

Create a reusable specification for the run: the held-out target, permitted context, respondents, model, and primary metric. The tutorial includes all 100 excerpt respondents, so no sampling or outcome-based stratification is needed.

zwill twin-experiment init-plan \
  --survey pew_w158 \
  --plan-id carbon_tax_v1 \
  --heldout-questions carbon_tax \
  --context-questions tree_planting,carbon_capture_credit,zero_carbon_power,methane_leaks,home_efficiency_credit \
  --model openai:gpt-5.5 --primary-metric nll \
  --path carbon_tax_v1.json

The response reports prediction_count_estimate: 100 and writes the plan beneath zwill’s output root as zwill_work/carbon_tax_v1.json. Validation checks question names, context policy, arms, and the prediction-count calculation without approving or running anything:

zwill twin-experiment validate \
  --input-path zwill_work/carbon_tax_v1.json

Expected result: status: ok, no validation errors, one held-out question, one arm, one model, and an estimated 100 predictions. Treat any error or unexpected count as a stop condition.

Read the plan itself before approving it. Confirm that carbon_tax is the only held-out target, precisely five different questions are context, all 100 eligible respondents are included, the model is intended, and the estimate is 100 calls.

zwill twin-experiment approve \
  --input-path zwill_work/carbon_tax_v1.json \
  --approved-by tutorial_reader \
  --note "Approved one target, five context questions, all 100 eligible respondents, one model, and NLL as the primary metric."

Approval records who reviewed which design and when. It is evidence of a deliberate specification, not a claim that the design is scientifically sufficient.

Expected result: the response shows approval.approved: true, approved_by: tutorial_reader, the review note, and the unchanged 100-prediction estimate. The command updates the plan file in place, so export-plan below reads the approved version.

Export and inspect the twin job

zwill twin-experiment export-plan \
  --input-path zwill_work/carbon_tax_v1.json \
  --output-dir carbon_tax_v1_jobs

ls carbon_tax_v1_jobs

The directory contains manifest.json plus an EDSL Jobs package whose generated name begins 01_baseline_ and ends _jobs.ep. The manifest maps that path to the approved arm and records the expected 100 predictions. Substitute the actual filename below:

ep inspect carbon_tax_v1_jobs/<jobs.ep>

Expected inspection: one EDSL probability question, 100 scenarios, one model, and therefore 100 runnable jobs. If those counts differ, stop and fix the plan rather than running an unexpectedly large or incorrectly scoped package.

Run with EDSL and import

ep run carbon_tax_v1_jobs/<jobs.ep> \
  --output carbon_tax_twin_results.ep

zwill twin-results import --survey pew_w158 \
  --input-path carbon_tax_twin_results.ep

The completed tutorial run produced 100 of 100 interviews with no exceptions. Import extracted all 100 predictions with zero issues, returned job_id: f87aa97344addd4a, and reported a total model cost of $1.0234. Your run receives its own ID; copy it into <twin_job_id>. Investigate and retry malformed rows instead of silently analyzing a selective subset.

08 Audit the prompts before scoring

Audit at least one imported run before looking at its performance. This order matters: inspect the construction before a promising score can make you more willing to excuse leakage or an unintended prompt. The run report exposes construction metadata, prompt templates, rendered scenarios, raw responses, malformed rows, and import issues.

zwill twin-results run-report --survey pew_w158 \
  --job-id <twin_job_id> --format html --path twin-audit.html

Open zwill_work/twin-audit.html and inspect what was actually sent to the model. In particular, confirm that each scenario contains the five approved policy answers but not that respondent’s carbon_tax answer, the weighted target marginal, or any other target-revealing material.

The command response confirms the written HTML path and audited job ID. The page should report 100 imported rows, zero import issues, the approved construction metadata, and rendered examples containing five context answers. Record any discrepancy rather than relying only on the score tables.

Leakage rule. The held-out answer, target-revealing follow-ups, downstream variables, and target summaries must be absent from twin prompts. A leaky prediction is not evidence of a capable twin.

This manual prompt inspection checks the study’s construction. The statistical leakage audit in the validation bundle provides a complementary check for unusually strong context–target relationships. Both must be clean before performance comparisons are credible.

09 Score held-out answers

With the construction audit complete, inspect which artifact was imported and then score its person-specific probabilities against the answers retained locally.

Inventory the imported run

zwill twin-study list --survey pew_w158

The returned table should show one imported job, 100 rows, zero issues, carbon_tax as the held-out question, and the model you ran. This is the quickest check that zwill is about to score the intended artifact.

job_id            rows  issues  heldout_questions  models
f87aa97344addd4a  100   0       carbon_tax         openai:gpt-5.5

Read the respondent-level twin report

zwill twin-results report --survey pew_w158 \
  --job-id <twin_job_id>

The command returns two tables with different purposes. The first is a case-by-case twin report: one row for each respondent’s hidden answer.

                              pew_w158 digital twin report
┏━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━┳━━━━━━━━┳━━━━━━┓
┃ respondent ┃ heldout    ┃ actual ┃ model          ┃ p(actual) ┃ uniform ┃ empirical ┃ nll   ┃ brier  ┃ top1 ┃
┡━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━╇━━━━━━┩
│ …100637    │ carbon_tax │ Oppose │ openai:gpt-5.5 │ 0.820     │ 0.500   │ 0.392     │ 0.198 │ 0.065  │ 1    │
│ …108592    │ carbon_tax │ Favor  │ openai:gpt-5.5 │ 0.900     │ 0.500   │ 0.608     │ 0.105 │ 0.020  │ 1    │
│ …          │ …          │ …      │ …              │ …         │ …       │ …         │ …     │ …      │ …    │
└────────────┴────────────┴────────┴────────────────┴───────────┴─────────┴───────────┴───────┴────────┴──────┘

What these two respondents show

Respondent …100637 opposed the carbon tax. Their other answers were mixed but mostly skeptical of the related policy proposals: they opposed the carbon-capture credit, zero-carbon power requirement, and home-efficiency credit, while favoring tree planting and methane-leak rules. The twin assigned Oppose probability 0.82. That is substantially more informative than the empirical marginal, which assigns only 0.392 to Oppose. The resulting NLL is just -log(0.82) = 0.198, the Brier score is 2 × (1 - 0.82)² = 0.065, and top-1 is correct.

Respondent …108592 favored all five context policies and also favored the held-out carbon tax. The twin assigned the observed answer probability 0.90, compared with 0.608 from the weighted empirical marginal. Its NLL is therefore 0.105 and its Brier score is 0.020. This is the intuitive behavior we want from a person-specific model: move away from the population average when that respondent’s permitted answers provide relevant evidence.

These are successful cases, not proof by anecdote. A confidently wrong row would show the reverse pattern: low p(actual), high NLL and Brier, and top1 = 0. Inspect those failures in the HTML or CSV report as carefully as these successes. The aggregate table and paired bootstrap comparison determine whether the pattern survives across all 100 respondents.
ColumnWhat it captures for one respondent
respondentThe deidentified join key used to pair prediction and observed truth.
heldoutThe question hidden from the model—carbon_tax on every row here.
actualThe answer retained locally and revealed only after inference for scoring.
modelThe model that returned the probability distribution.
p(actual)The probability assigned to the answer that person actually gave. Higher is better and partial credit is possible.
uniformThe probability from splitting mass equally over legal options: 0.5 for either answer here.
empiricalThe committed weighted marginal probability of this person’s actual option. It uses population truth, not person-specific evidence, so it is an oracle reference rather than a deployable predictor.
nll-log(p(actual)). Lower is better; confidently wrong predictions receive a large penalty.
brierSquared error across both Favor and Oppose probabilities. Lower is better.
top11 when the higher-probability option equals the actual answer, otherwise 0.

Read the aggregate model summary

The second table averages those case-level scores over all usable respondents. It answers how the run performed overall, while the first table shows which people generated that average.

                              model summary
┏━━━━━━━━━━━━━━━━┳━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┳━━━━━━┓
┃ model          ┃ rows ┃ p(actual) ┃ uniform p ┃ empirical p ┃ nll   ┃ uniform nll ┃ empirical nll ┃ brier ┃ uniform brier ┃ empirical brier ┃ top1 ┃
┡━━━━━━━━━━━━━━━━╇━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━╇━━━━━━┩
│ openai:gpt-5.5 │ 100  │ 0.690     │ 0.500     │ 0.523       │ 0.513 │ 0.693       │ 0.670         │ 0.329 │ 0.500         │ 0.477           │ 0.762│
└────────────────┴──────┴───────────┴───────────┴─────────────┴───────┴─────────────┴───────────────┴───────┴───────────────┴─────────────────┴──────┘
ColumnWhat the summary capturesHow to read it
rowsSuccessfully scored respondent–question predictions.It should be 100. Fewer rows require investigation.
p(actual)Mean probability assigned to each respondent’s observed answer.Higher is better; 0.5 is the binary uniform floor.
uniform pA no-information prediction split equally across legal options.Two options imply 1 / 2 = 0.5.
empirical pMean probability assigned by the weighted observed target marginal.An oracle aggregate reference: useful for comparison, unavailable for a genuinely new target.
nllMean negative log likelihood.Lower is better and confident errors are strongly penalized.
uniform nllNLL from assigning 0.5 to each option.-log(0.5) = 0.693; the twin should be below it.
empirical nllNLL from assigning everyone the committed weighted target distribution.Tests whether person-specific probabilities add signal beyond knowing the aggregate split.
brierMean squared error of the complete binary distribution.Lower is better.
uniform brier / empirical brierThe same distributional error for the two reference predictors.The twin should be lower, but the XGBoost comparison remains the stronger deployable test.
top1Share whose highest-probability option matches the observed answer.Higher is better, but it ignores confidence.

These are the weighted results from the completed tutorial run. The twin beat uniform and the weighted empirical-marginal oracle on mean NLL: 0.513 versus 0.693 and 0.670 respectively. Its weighted top-1 accuracy was 76.2%. Those comparisons are encouraging, but XGBoost is the more demanding respondent-conditioned benchmark.

XGBoost is not in this first report yet. This command selected only the imported language-model job. Uniform is an automatic no-information reference. The next chapter fits XGBoost, appends its predictions as a second model, and creates the paired comparison.

Write reusable reports and aggregate diagnostics

zwill twin-results report --survey pew_w158 \
  --job-id <twin_job_id> --format html --view full \
  --path twin-results.html \
  --calibration-svg-path twin-calibration.svg

zwill twin-results report --survey pew_w158 \
  --job-id <twin_job_id> --format csv \
  --path twin-results.csv

zwill twin-results marginal-diagnostics --survey pew_w158 \
  --job-id <twin_job_id> \
  --target probability-job \
  --target-job-id <probability_job_id> \
  --format svg --path twin-vs-one-shot-marginals.svg

The HTML supports visual inspection, the CSV provides one scored row per respondent, and twin-calibration.svg is a reusable reliability chart: points on the diagonal are calibrated, while points below it indicate overconfidence. Marginal diagnostics compare the average of 100 person-specific distributions with the one-shot population prediction; the SVG makes the option-by-option gaps visible. Those aggregate comparisons complement—but cannot replace—the respondent-level scores.

What marginal diagnostics returns: the selected twin and one-shot job IDs, aligned option labels, the mean twin-implied distribution, the one-shot distribution, and distance measures between them. A small distance means the two methods agree about the aggregate; it says nothing by itself about matching individuals.

10 Compare the twins with XGBoost

Now ask whether the language-model twins add predictive value beyond a cheaper conventional model. This is a separate question from prompt auditing: a perfectly non-leaky twin can still perform no better than an ordinary supervised baseline.

Fit the conditional XGBoost baseline

Beating uniform probabilities is a weak test. A useful digital twin should also beat a conventional supervised learner given comparable respondent information. twin-validate therefore fits zwill’s conditional XGBoost baseline. It embeds the question texts and each question–option pair, represents each respondent by their non-held-out selections, and learns how those semantic features predict an option.

The fit is leave-one-question-out across survey items. For carbon_tax, XGBoost trains on the 100 respondents’ labeled rows from the other five questions—1,000 candidate-option training rows for five binary items. The target question’s own answers and marginal never enter training. It then transfers what it learned across questions to predict Favor/Oppose for the same 100 people scored by the language-model job, making the comparison paired.

What kind of XGBoost baseline is this? This is a semantic cross-question baseline, not ordinary cross-validation trained on some respondents’ carbon_tax labels. That stricter exclusion makes it deployable to a genuinely new target, but it answers a different question from “how well can supervised learning predict this target once labeled target examples exist?” For that latter use case, add a respondent-fold baseline trained on target labels and keep train/test respondents separate. Zwill can also include readable respondent covariates; this excerpt deliberately contains none because the demographic codebook is not bundled.
zwill twin-validate --survey pew_w158 \
  --jobs <twin_job_id> --out report_out/validation \
  --embedder edsl --require-baseline \
  --n-boot 2000 --seed 42

--embedder edsl reproduces this tutorial run by obtaining text-embedding-3-small vectors through Expected Parrot. --require-baseline makes validation fail rather than quietly omit XGBoost if embedding or fitting cannot complete. For an API-free run, use --embedder sentence-transformers; changing the embeddings can change the baseline and must be recorded as a different specification.

On success, the baseline step reports model_type: xgboost, training_rows: 1000, prediction_rows: 100, and scored_questions: ["carbon_tax"]. Zwill stores those predictions under baseline:conditional-embedding and computes exactly the same NLL, Brier, p(actual), and top-1 fields used for the language-model twins.

Read the twin-versus-XGBoost result

Representative completion output: status: ok, baseline.model_type: xgboost, baseline.training_rows: 1000, baseline.prediction_rows: 100, bootstrap.resamples: 2000, and paths for report.html, report_data.json, bootstrap.json, bootstrap-intervals.svg, calibration.svg, leakage_audit.json, and manifest.json. A missing required baseline or failed leakage step makes the command fail instead of yielding a partial success.

Open zwill_work/report_out/validation/report.html. Unlike the earlier one-job report, its model summary contains both the language-model twin and baseline:conditional-embedding. The respondent-level table also includes XGBoost rows, so you can see whether both methods fail on the same people or in different ways.

The validation command also resamples respondents to estimate paired bootstrap intervals for the difference between each twin model and XGBoost. Those results are in zwill_work/report_out/validation/bootstrap.json. For p(actual) and top-1, higher is better; for NLL and Brier, lower is better. Interpret the direction of each twin-minus-baseline difference accordingly, and do not claim an improvement when its uncertainty interval includes zero.

Results from the completed tutorial run

MetricGPT-5.5 twinXGBoostPaired twin − XGBoost [95% CI]Interpretation
p(actual)0.6900.671+0.019 [−0.033, +0.075]Point estimate favors the twin; interval includes zero.
NLL0.5130.698−0.186 [−0.370, −0.034]Twin is better; the full interval is below zero.
Brier0.3290.420−0.091 [−0.189, +0.002]Point estimate favors the twin; interval narrowly includes zero.
Top-1 accuracy0.7620.764−0.002 [−0.135, +0.118]Essentially tied and statistically inconclusive.
Forest plot of paired bootstrap differences between GPT-5.5 twins and conditional XGBoost
This is the same forest-plot form that zwill writes to zwill_work/report_out/validation/bootstrap-intervals.svg. Each dot is the paired twin-minus-XGBoost estimate and each line is its 95% bootstrap interval. Only NLL clears zero in this run.

NLL was declared as the primary metric before export. On that metric, the twin produced meaningfully better probabilities than XGBoost: its paired interval clears zero in the favorable, negative direction. The methods had nearly identical top-1 accuracy, so the gain came from probability quality rather than getting more binary choices correct. The leakage audit examined all five context–target pairs and flagged none.

ArtifactWhat it captures
zwill_work/report_out/validation/report.htmlThe respondent-level and aggregate scores for both the language-model twins and the conditional XGBoost baseline.
bootstrap.jsonPoint estimates and paired confidence intervals for twin-versus-XGBoost metric differences.
bootstrap-intervals.svgA publication-ready forest plot of those paired differences and their uncertainty intervals.
calibration.svgA standalone reliability plot showing whether stated confidence matches observed accuracy.
leakage_audit.jsonChecks for suspicious relationships that could make the held-out answer recoverable from supposedly permitted inputs.
report_data.jsonThe structured inputs behind the validation report, including twin rows, baseline rows, bootstrap results, and audit summary.
manifest.jsonThe job IDs, settings, steps, and paths needed to trace and reproduce the validation bundle.
Interpret this as a pilot. Unlike the former six-person fixture, this study really fits XGBoost and produces a paired 100-person comparison. It is still only one target from one related question battery. Treat the result as directional until it survives more held-out questions, a larger sample, and prespecified model or prompt comparisons.
zwill next --survey pew_w158

Run zwill next after validation and follow any remaining audit, sample-size, or approval steps it reports before treating the evidence as complete.

11 Read the evidence

Run zwill guide show interpreting-results before writing conclusions. Lower NLL and Brier scores are better. Accuracy is useful as a sanity check, but the headline is the paired difference against the conditional baseline.

Evidence
Supports a claim
Does not support it
Leakage audit
Clean
Target information exposed
Twin − conditional baseline
Improvement
Tie or deterioration
Bootstrap interval
Clears zero
Includes zero
Calibration / ECE
Probabilities track outcomes
Overconfident misses

For this run, the NLL metric prespecified in the approved plan passes the table’s comparison gate: −0.186 [−0.370, −0.034] against XGBoost. This is not a public preregistration: the tutorial target was chosen while developing the example. The Brier interval barely crosses zero, while p(actual) and top-1 remain inconclusive. That is mixed but useful evidence: the twin assigned probabilities better on the primary scoring rule without improving classification accuracy.

Also inspect calibration, individual failures, malformed-row rate, and marginal fit. Expected calibration error (ECE) groups predictions by confidence and compares average confidence with observed frequency; lower is better. The twin’s ECE in this run was 0.0446, but a single small binary sample can make calibration estimates unstable, so inspect the calibration plot and confident misses as well. The empirical marginal is an oracle benchmark because the target distribution has already been observed; it is not deployable for a genuinely new target.

Appropriate conclusion for this pilot. “For this 100-person Pew W158 excerpt and the held-out carbon-tax item, GPT-5.5 improved NLL relative to the leave-one-question-out XGBoost baseline, while accuracy was indistinguishable.” Do not generalize from one climate-policy item to digital-twin validity overall.

12 Assemble the evidence and author the report

Zwill’s responsibility ends with a rich, auditable evidence package. It should calculate the metrics, retain respondent-level predictions and raw model evidence, document the study design, fit the baselines, run the leakage and bootstrap analyses, and expose all of that in structured files. It should not decide the final narrative or send a second EDSL job to a model to write one.

Build the evidence package

zwill report build --survey pew_w158 --path report_out \
  --job-id <twin_job_id> \
  --probability-job-id <probability_job_id>

This command assembles the survey profile, one-shot results, twin scores, run audit, and available diagnostics into zwill_work/report_out/. Because validation was written to zwill_work/report_out/validation/, zwill detects that bundle and links its XGBoost comparison, bootstrap intervals, leakage audit, and structured evidence from the consolidated index. The result is one publishable evidence tree rather than two disconnected folders.

Top of a generated zwill evidence-package index showing the survey overview and available report sections
The generated zwill_work/report_out/index.html is the front door to the evidence package. It inventories the survey and links the profile, one-shot analysis, twin validation, diagnostics, and prompt audit instead of presenting a single unsupported conclusion.
Top of a generated zwill twin-validation report showing run overview and model-performance tables
A representative Pew W158 validation page generated locally by zwill. The exact rows in this screenshot come from another completed W158 run and illustrate the package layout; the numerical walkthrough in this tutorial continues to use the 100-person carbon_tax run reported above.

What the structured evidence looks like

The HTML is for inspection; JSON is what lets a coding agent trace claims without scraping a page. For example, a compact summary of this tutorial’s completed run looks like this. The complete package keeps more detail, including respondent rows, calibration bins, every bootstrap comparison, and the audit pairs.

{
  "survey": "pew_w158",
  "heldout_question": "carbon_tax",
  "respondents": 100,
  "twin": {
    "model": "openai:gpt-5.5",
    "negative_log_likelihood": 0.5129,
    "top1_accuracy": 0.7621,
    "expected_calibration_error": 0.0446
  },
  "conditional_baseline": {
    "model_type": "xgboost",
    "negative_log_likelihood": 0.6984
  },
  "primary_comparison": {
    "metric": "negative_log_likelihood",
    "twin_minus_xgboost": -0.1855,
    "confidence_interval_95": [-0.3695, -0.0343]
  },
  "leakage_audit": {"pair_count": 5, "flagged_count": 0}
}

Open the complete tutorial evidence-summary JSON. An agent should verify the primary comparison’s direction and interval, confirm that the leakage audit is clean, and then use the respondent-level and calibration files to look for important failures before drafting prose.

Evidence for the coding agentWhat it contributes
zwill_work/report_out/index.htmlA navigable inventory of the survey profile and currently available technical pages.
zwill_work/report_out/validation/report_data.jsonStructured study design, twin and XGBoost rows, score summaries, bootstrap comparisons, calibration diagnostics, and audit results.
zwill_work/report_out/validation/report.htmlThe human-readable technical validation report against which narrative claims can be checked.
zwill_work/report_out/validation/bootstrap.jsonThe paired uncertainty intervals needed to distinguish evidence from sampling noise.
zwill_work/report_out/validation/leakage_audit.jsonThe context–target checks that must be read before making a positive performance claim.
zwill_work/twin-audit.htmlRendered prompts, respondent context, raw responses, construction metadata, and import issues.
examples/pew_w158_tutorial/source.mdSurvey provenance, scope, and the limitations of the bundled excerpt.

Have the coding agent write the final narrative

The coding agent should read those artifacts, decide which facts matter for the intended audience, and author the final HTML or Markdown report directly. That is a reasoning and communication task: it requires reconciling the clean leakage audit, the significant NLL advantage over XGBoost, the inconclusive accuracy comparison, the one-question scope, and the intended decision use. A generic generated paragraph cannot make those judgments reliably.

Ask the agent to keep every substantive claim traceable to a named artifact or reported value. The final report should explain the survey and holdout design, translate the metrics into plain language, compare the twins with XGBoost, discuss individual failures and calibration, state what the pilot supports, and distinguish that from what would require a larger multi-question study.

The division of labor. Zwill produces contextualized evidence and reproducible diagnostics. The coding agent studies that evidence and crafts the report. EDSL runs the survey inference jobs; it is not used for a separate report-writing round trip.

Publish the agent-authored report together with—or linked to—the technical evidence package. Readers should be able to move from the conclusion back to the validation table, bootstrap interval, leakage result, and underlying study design.

13 Scale from one target to a real study

The carbon-tax example teaches the mechanics, but one question cannot establish that twins work across a survey. A serious validation should hold out several questions chosen before inference. Select items that span different content, base rates, and levels of difficulty; record every eligible item considered and explain exclusions.

zwill twin-experiment init-plan --survey pew_w158 \
  --plan-id climate_policy_v2 \
  --heldout-questions tree_planting,carbon_tax,carbon_capture_credit,zero_carbon_power,methane_leaks \
  --sample-respondents 100 --seed 20260721 --stratify-actual \
  --context-question-count 5 \
  --model openai:gpt-5.5 --primary-metric nll \
  --path climate_policy_v2.json

zwill twin-experiment validate --input-path climate_policy_v2.json
zwill twin-experiment approve --input-path climate_policy_v2.json
zwill twin-experiment export-plan --input-path climate_policy_v2.json \
  --output-dir climate_policy_v2_jobs

init-plan creates an editable specification, validate resolves the question names and prediction count, and approve records that the design was reviewed before model results were visible. Export then creates the .ep jobs and a manifest. Inspect the manifest and confirm that its scenario count still matches the approved count before running anything. Open the example draft plan JSON.

respondents × held-out questions × approaches × models100 × 5 × 1 × 1 = 500 predictions in this teaching excerpt

Multi-target validation unlocks evidence the one-question pilot cannot provide: per-question forest plots, marginal rank-order accuracy, cross-question attenuation, joint-structure checks, and a leakage matrix across every context–target pair. Report heterogeneity rather than hiding it in one grand mean—a model that works on four easy policy items and fails on a fifth has taught you something important.

Sampling rule. Use the same respondent sample for every experimental arm. If filtering changes the eligible count or class balance after approval, stop and update the plan rather than silently changing the estimand.

14 Compare models and prompt constructions

Zwill treats model choice and prompt design as experiments, not as invisible tuning. A reusable approach records how the twin is constructed. An experiment plan crosses those approaches with models on the same respondents and held-out targets.

Define two prompt approaches

zwill twin-approach scaffold --survey pew_w158 --approach-id survey_only \
  --path approaches/survey_only.json
zwill twin-approach scaffold --survey pew_w158 --approach-id survey_plus_notes \
  --include-agent-material --tag context-ablation \
  --path approaches/survey_plus_notes.json

zwill twin-approach add --survey pew_w158 \
  --input-path zwill_work/approaches/survey_only.json
zwill twin-approach add --survey pew_w158 \
  --input-path zwill_work/approaches/survey_plus_notes.json
zwill twin-approach diff --survey pew_w158 \
  --left survey_only --right survey_plus_notes

Reusable approaches vary construction inputs such as context selection, metadata, supplemental material, and models. The example below is an ablation for a richer study that has first imported respondent notes with agent-material add; on the teaching excerpt, which has no such notes, the arms would be identical. Prompt pipelines are a complementary experiment: a pipeline such as examples/twin_pipelines/dialectical.json is an ordered set of EDSL questions. An early step can weigh evidence, and the final step returns the scoreable probability distribution. Zwill preserves intermediate reasoning for audit but scores only the final answer.

zwill edsl build --survey pew_w158 --target twin-probability-job \
  --heldout-question carbon_tax --sample-respondents 100 --seed 42 \
  --twin-prompt-pipeline examples/twin_pipelines/dialectical.json \
  --model openai:gpt-5.5 --allow-unapproved \
  --path dialectical_jobs.ep

# After running and importing matched raw and pipeline jobs:
zwill twin-experiment record --survey pew_w158 \
  --experiment-id carbon_prompt_ab --job-id <raw_job_id> \
  --approach raw --primary-metric nll
zwill twin-experiment record --survey pew_w158 \
  --experiment-id carbon_prompt_ab --job-id <pipeline_job_id> \
  --approach dialectical --primary-metric nll

For a real comparison, put reusable construction approaches and every model in the approved plan. Record custom pipeline jobs under one experiment ID after import. In both cases, keep the sample and seed fixed so differences are paired.

{
  "defaults": {
    "sample_respondents": 100,
    "seed": 20260721,
    "model": ["openai:gpt-5.5", "google:gemini-2.5-pro"]
  },
  "arms": [
    {"approach_id": "survey_only"},
    {"approach_id": "survey_plus_notes"}
  ]
}

Those fields are the relevant excerpt from the plan, not a second complete plan file. With five targets, the design now requests 100 × 5 × 2 × 2 = 2,000 predictions. Run validate again, inspect the revised count and estimated cost, and obtain a new approval before export.

zwill twin-experiment bundle --survey pew_w158 \
  --plan-id climate_policy_v2 --metric nll \
  --output-dir report_out/experiment

zwill twin-experiment dashboard --survey pew_w158 \
  --plan-id climate_policy_v2 --path report_out/experiment/dashboard.html
zwill twin-experiment compare --survey pew_w158 \
  --experiment-id climate_policy_v2 \
  --metric nll
zwill twin-experiment microdata --survey pew_w158 \
  --experiment-id climate_policy_v2 \
  --path report_out/experiment/microdata.html

The comparison ranks arms on the declared metric; the dashboard summarizes them; the microdata page reveals which respondents each approach helped or harmed. A reasoning prompt earns its extra calls only if it improves held-out scores or calibration—not because its prose sounds thoughtful.

15 Use the full diagnostic package

The validation report answers whether a particular run beats its baselines. The broader evidence package asks additional questions: Does the twin preserve question ordering? Does it add individual signal? Are gains concentrated among a few respondents? Does it reproduce relationships among answers?

zwill report facts --survey pew_w158 --path report_out
zwill report analyze --survey pew_w158 --path report_out \
  --job-id <twin_job_id> --probability-job-id <probability_job_id>
zwill report render --survey pew_w158 --path report_out \
  --job-id <twin_job_id> --probability-job-id <probability_job_id>
DiagnosticQuestion it answersHow to read it
Actual-answer lift histogramWho gained probability on the answer they actually gave?Mass to the right of zero indicates individual-level improvement; inspect the left tail for harmed respondents.
Empirical-marginal lift histogramDoes personalization beat simply assigning everyone the observed population split?A positive center is evidence of person-specific signal; the empirical benchmark remains an oracle.
Marginal rank orderAcross questions, does the model know which options or items are more prevalent?Useful aggregate structure does not by itself validate individuals.
Pairwise-order accuracyHow often are pairs of aggregate quantities ordered correctly?Read the permutation result beside the point estimate to distinguish skill from chance.
Conditional consistencyDo predictions change coherently across respondent groups and contexts?Large unexplained gaps can reveal brittle prompts or missing covariates.
Joint structureAre relationships among questions preserved rather than washed out?Attenuated correlations mean plausible marginals may still conceal unrealistic synthetic people.
Subgroup marginalsWhere does performance or prevalence differ?Require adequate subgroup counts and report uncertainty; do not hunt for tiny favorable slices.
Histogram of probability lift on respondents' actual answers relative to a uniform baseline
Actual-answer lift. Bars show how much probability the twin gained or lost on the answer each person actually gave, relative to uniform. Look at the full distribution, not only its mean.
Histogram of respondent-level probability lift relative to the empirical marginal
Lift over the empirical marginal. This harder oracle comparison asks whether personalization adds signal beyond assigning everyone the population split.
Pairwise aggregate-order accuracy chart with a chance reference
Aggregate ordering. This chart asks whether predicted aggregate quantities have the right relative order. These three plots come from the representative locally generated W158 package shown above; they demonstrate zwill’s output form rather than adding results to the tutorial’s separate 100-person carbon-tax run.

These artifacts appear in both human-readable pages and machine-readable files under zwill_work/report_out/analysis/ and zwill_work/report_out/data/. The report catalog and manifests tell an agent which files are ready and where each value came from.

Optional marginal calibration

zwill twin-results calibrate-marginal --survey pew_w158 \
  --job-id <twin_job_id> --questions carbon_tax \
  --target probability-job --target-job-id <probability_job_id> \
  --output-job-id twin_calibrated_to_one_shot

zwill twin-results compare-report --survey pew_w158 \
  --jobs <twin_job_id>,twin_calibrated_to_one_shot \
  --format html --path report_out/calibration-comparison.html

This adjusts person-level distributions so their aggregate matches the chosen target while retaining as much individual ordering as possible. Calibration to a one-shot forecast can be deployable if that forecast would truly be available. Calibration to the observed empirical marginal is an oracle sensitivity analysis and must never be presented as out-of-sample performance.

16 Add covariates and supplemental evidence responsibly

The teaching excerpt deliberately omits demographics. In a real panel, readable respondent metadata—such as age group, region, or party identification—can be legitimate predictive context for both the twin and XGBoost. Zwill includes respondent metadata by default and renders it as a separate profile block.

{"respondent_id":"r_001","metadata":{
  "age_group":"50–64",
  "region":"Midwest",
  "party_identification":"Independent"
}}

Expand source codes before import: "age_group":"50–64" is meaningful evidence, while "F_AGECAT":4 is not. Use exclusions when a field is an identifier, downstream outcome, proxy for the held-out answer, or inappropriate for the intended deployment.

zwill edsl build --survey <survey> --target twin-probability-job \
  --heldout-questions <q1,q2> \
  --exclude-metadata-key panelist_id \
  --exclude-metadata-key target_followup \
  --model openai:gpt-5.5 --approved-plan study.json \
  --path twin_jobs.ep

For richer qualitative evidence, agent-material and --twin-material can attach approved interview notes or other respondent-specific context. Cap its length, retain its provenance, audit rendered prompts, and run an ablation against a survey-only approach. More context is not automatically better.

Reuse respondents as an EDSL AgentList

Digital-twin validation asks whether agents predict answers that are already known. After that validation, zwill can also export the constructed respondents as a durable EDSL AgentList.ep and ask them a genuinely new question. Keep this prospective exercise separate from the held-out validation evidence.

zwill edsl build --survey pew_w158 --target agent-list \
  --path pew_w158_agents.ep
zwill agent-list inspect --input-path pew_w158_agents.ep

zwill agent-study export --agent-list pew_w158_agents.ep \
  --question-name new_policy --question-type multiple_choice \
  --question-text "Would you favor this new proposal?" \
  --question-option Favor --question-option Oppose \
  --model openai:gpt-5.5 --path new_policy_jobs.ep
ep run new_policy_jobs.ep --output new_policy_results.ep
zwill agent-study import --input-path new_policy_results.ep
zwill agent-study report --format html --path report_out/new-policy.html

The resulting forecast is an application of the validated construction, not additional validation: there is no observed answer for scoring. State that distinction prominently whenever prospective agent-study results are published.

Fairness and leakage boundary. A covariate may improve prediction and still be inappropriate to use. Document why every sensitive or proxy field is necessary, measure subgroup behavior, and remove anything that would be unavailable or disallowed at deployment.

17 Recover failures and deliver reproducible artifacts

Retry only malformed model rows

Provider hiccups should not silently shrink the scored sample or force an expensive full rerun. If import reports malformed rows, build a retry job from the original package, run it, and merge only the recovered scenarios.

zwill twin-results retry-malformed --survey pew_w158 \
  --job-id <twin_job_id> --job twin_jobs.ep \
  --path retry_jobs.ep
ep run retry_jobs.ep --output retry_results.ep
zwill twin-results import --survey pew_w158 \
  --job-id <twin_job_id> --merge \
  --input-path retry_results.ep

Resolve ingestion problems explicitly

zwill quarantine list --survey pew_w158
zwill quarantine resolve --survey pew_w158 --issue-id <issue_id> \
  --action resolve \
  --note "Mapped source code 4 to 50–64 from the published codebook"
zwill commit --survey pew_w158

Quarantine keeps unknown codes, invalid options, and incomplete records visible. Resolve the source problem and record what changed; do not simply suppress an issue to reach the model-running stage.

Export, package, and automate

zwill twin-results export --survey pew_w158 --job-id <twin_job_id> \
  --format long --path predictions-long.csv
zwill twin-results export --survey pew_w158 --job-id <twin_job_id> \
  --format wide --path predictions-wide.csv
zwill twin-results package --survey pew_w158 --job-id <twin_job_id> \
  --path predictions.csv --zip-path predictions.zip

zwill workflow dry-run study-workflow.json
zwill workflow explain study-workflow.json
zwill workflow run study-workflow.json

Long CSV is best for analysis; wide CSV is convenient for respondent-level joins; the ZIP is a portable delivery artifact. Declarative workflows make repeated builds inspectable: always dry-run and explain them before execution. For very large true-holdout studies, twin-study export-holdout creates chunked jobs so failures and reruns stay bounded.

18 Validate numeric, ranking, and open-ended targets

The main gate handles closed-ended choices. Zwill routes other question types through scoring rules appropriate to what the model predicts.

Numeric questions: predict quantiles

zwill edsl build --survey <survey> --target numeric-twin-job \
  --heldout-question household_income --allow-unapproved \
  --path numeric_jobs.ep
ep run numeric_jobs.ep --output numeric_results.ep
zwill numeric-results import --survey <survey> \
  --input-path numeric_results.ep
zwill numeric-results report --survey <survey> --job-id <job_id> \
  --format html --path report_out/numeric.html

The twin returns p05, p25, p50, p75, and p95 rather than category probabilities. Read pinball loss, CRPS, and interval coverage against a marginal-quantile climatology baseline.

Numeric metricMeaningBetter
Pinball lossScores each reported quantile, penalizing misses differently above and below it.Lower
CRPSScores the quality of the predicted distribution across the outcome scale.Lower
Median MAEAbsolute error of the predicted p50 as a point forecast.Lower
50% / 90% coverageShare of outcomes inside the corresponding predicted interval.Close to 50% / 90%
Pinball skillImprovement over giving everyone the population marginal quantiles.Positive

Ranking and MaxDiff: predict utility

zwill edsl build --survey <survey> --target rank-utility-twin-job \
  --rank-task-id priorities --allow-unapproved --path rank_jobs.ep
ep run rank_jobs.ep --output rank_results.ep
zwill twin-results import --survey <survey> --input-path rank_results.ep
zwill twin-results rank-report --survey <survey> \
  --rank-task-id priorities --format html \
  --path report_out/rank-priorities.html

Declare a stable rank_task_id during import. Rank tasks have their own leakage and partial-ranking rules and are intentionally not folded into the multiple-choice headline gate.

Rank metricMeaningBetter
SpearmanCorrelation between predicted and actual ordering among ranked items.Toward 1
PairwiseShare of item pairs placed in the correct relative order.Above 0.5 chance
Top-K identificationWhether the predicted top-K catches the respondent’s actual selected set—the headline for partial rankings.Above K/N chance
Rank MAEAverage absolute error in rank position.Toward 0
Top-1Share whose predicted first item matches the actual first item.Higher

Open text: derive themes, then validate the coded target

zwill edsl build --survey <survey> --target open-codebook-job \
  --heldout-question climate_concern_text --n-themes 8 \
  --model openai:gpt-5.5 --path codebook_jobs.ep
ep run codebook_jobs.ep --output codebook_results.ep
zwill open-coding codebook-import --survey <survey> \
  --input-path codebook_results.ep

zwill edsl build --survey <survey> --target open-coding-job \
  --heldout-question climate_concern_text --model openai:gpt-5.5 \
  --path coding_jobs.ep
ep run coding_jobs.ep --output coding_results.ep
zwill open-coding import --survey <survey> \
  --input-path coding_results.ep \
  --coded-question-name climate_concern_theme

The first job proposes a codebook; the second applies one theme consistently to every answer. Inspect theme coverage and the unclassified rate, then validate climate_concern_theme as an ordinary multiple-choice holdout. The coding model and twin model solve different tasks, so preserve both artifacts.

Open-text artifactWhat to inspect
Codebook resultWhether themes are distinct, exhaustive enough, interpretable, and supported by sampled answers.
Coded rowsRaw answer, assigned theme, model output, and any malformed or unclassified cases.
Coding reportTheme frequencies and unclassified share; above 20% warns that the codebook may not fit.
Held-out validationWhether twins predict the coded theme against the same conditional baseline and leakage gate used for other categorical targets.

19 From tutorial to a defensible study

  1. Use the complete documented survey rather than the 100-person teaching excerpt, and include readable codebook-expanded covariates when they are legitimate predictors.
  2. Select 5–10 defensible held-out questions and declare leakage exclusions.
  3. Create and approve a twin-experiment plan with sample, seed, models, approaches, cost, outputs, and decision criteria.
  4. Prefer at least 200 scored respondents—or all eligible respondents when fewer exist—and flag sparse options and subgroups.
  5. Compare prompt or model variants through the same gate; do not choose a winner from a bare run.
# The two commands to keep using throughout the study:
zwill guide
zwill next --survey <survey_name>

zwill also supports numeric targets through quantile predictions, ranking and MaxDiff batteries through a separate rank-utility flow, and open-ended targets by first coding responses into themes. The bundled guide routes each question type to its proper validation workflow.