Evidence asset

Survey data dictionary template: fields, codes and skip logic

Document survey variables, response codes, missing values, derivations and branch rules with a reusable 15-field template and reproducible audit.

Published
11 September 2026
Reading time
11 min
Author and reviewer
The Survey Review

A survey data dictionary is the contract between the questionnaire and its analysis file. It gives every exported field a stable name, definition, type, allowed values, missing-value rule and routing note, so an analyst can interpret the data without guessing from the survey screen.

What is a survey data dictionary?

A survey data dictionary, also called a codebook, describes the structure, content and variable-level meaning of a survey dataset. Harvard Biomedical Data Management identifies variable names, readable names, units, allowed values and definitions as core elements. See Harvard's data-dictionary guidance.

For survey files, the dictionary should also retain question wording, response labels, missing-value meanings and the universe or skip pattern. ICPSR's current codebook guidance treats those as variable-level details needed to make a file self-explanatory. See ICPSR's codebook guidance.

What belongs in a survey data dictionary?

Use the 15-field template below as a minimum working specification. Add study-level methods, access rules, weights or frequency summaries when the analysis requires them.

FieldPurposeExample
variable_nameExact, stable export fieldsatisfaction_overall
question_idStable questionnaire referenceQ12
labelShort readable nameOverall satisfaction
definitionExact concept representedSatisfaction with service in the last 30 days
question_textWording respondents sawOverall, how satisfied were you?
typeStorage typeinteger
allowed_valuesValid stored codes1, 2, 3, 4, 5
value_labelsMeaning of each code1 Very dissatisfied; 5 Very satisfied
missing_valuesNon-substantive codes-7 Refused; -8 Not applicable; -9 Missing
unit_or_periodUnit and recall windowLast 30 days
validationRange, format or constraintInteger 1 through 5
branch_ruleUniverse in which the item appearsASK IF Q2 = 1
derivationFormula for calculated fieldsMean of Q12a through Q12d
sourceOrigin of item or constructQuestionnaire v1.3
steward_revisionOwner, version and change noteResearch Ops; v1.3; wording clarified

How should survey variables be named?

Choose one machine-readable convention and apply it consistently. This template uses lowercase snake case, begins names with a letter and keeps display order out of the variable name. A generated name such as q12 can become misleading when a question is inserted; satisfaction_overall survives that layout change. Keep the stable questionnaire ID separately to preserve lineage.

The exact convention is a team choice, not a universal standard. The non-negotiable rule is that dictionary names match the data file exactly. OSF's current guide also calls for variable names exactly as they appear in the spreadsheet, alongside readable names, units, allowed values and definitions. See the OSF template guidance.

How should missing survey answers be coded?

Do not let one blank stand for several events. Distinguish at least not shown by logic, shown but unanswered, refused, not applicable and technical failure when those states matter. Document both code and label, and configure analysis software so a numeric missing code is not treated as a real measurement.

Never reuse a variable name when the represented concept or scale changes materially. Create a new variable or version, record the mapping and keep the questionnaire version with the dictionary.

Worked example: branch logic and structural missingness

Question Q2 asks whether a respondent used the service. Only users see Q12 satisfaction. For a non-user, satisfaction_overall should carry a documented structural-missing value such as -8 Not applicable, not zero. Zero could be mistaken for the lowest score and corrupt a mean.

The dictionary row states ASK IF Q2 = 1, allowed range 1 through 5, -7 for refused and -8 for not applicable. A validation query can assert that every Q2 non-user has -8, while every Q2 user has 1 through 5, -7 or a separately documented technical-missing code.

Reproducible dictionary audit

  1. Export one synthetic response for every terminal survey route.
  2. Compare exported columns with dictionary variable names.
  3. Confirm every stored code has exactly one label and every documented label has a code.
  4. Test the minimum, maximum and one invalid value for each numeric field.
  5. Verify that skipped items differ from unanswered items that were displayed.
  6. Recalculate each derived variable from its source fields.
  7. Confirm every branch rule references a current, stable question ID.
  8. Store the dictionary, questionnaire and processing-code versions together.

How do wide, long and multi-select exports differ?

A wide file may create one column per answer option, while a long file may create one row per selected option. Document the shape and uniqueness key. For checkboxes, state whether an unselected option means false, not shown or missing; do not infer that meaning from an empty cell.

If the dataset will move between repositories or software, a structured standard can preserve more of this metadata. DDI Alliance says its codebook model can capture record layout, variable names and labels, categories, missing codes, universe statements and provenance in machine-actionable form. See DDI Alliance's codebook overview.

Limitations

A dictionary documents meaning but does not validate a measure, prove the sample represents a population, protect personal data or guarantee reproducible analysis. Pair it with the questionnaire, field disposition record, processing code and access controls. Remove direct identifiers from public examples. This template is intentionally software-neutral and does not test any vendor's export behavior.

Sources and limitations

Verification date: 11 September 2026. This is operational survey-design guidance, not legal advice. Requirements can differ by jurisdiction, audience and research purpose. Send corrections with a primary source through our corrections process.