Skip to content

Latest commit

 

History

137 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

formtransform

AI-Assisted

One library to transform surveys between the standards of the CDL survey ecosystem — XLSForm (Kobo Toolbox), LimeSurvey TSV, and DDI Codebook — built on a canonical survey type registry (registry/) that defines exactly what is supported and how the standards map onto each other.

What this repo ships

Four things other projects consume — all derived from the one registry:

Artifact Where Consumed by
@correlaid/formtransform — TypeScript library + formtransform CLI npm package (github:CorrelAid/formtransform) formtransform-app, qwacback, direct CLI use
schematron-worker image — Java NATS service validating DDI against the XSDs + CDL rules, baked in ghcr.io/correlaid/schematron-worker:<version>, built from workers/schematron-worker/ qwacback (runs the image, does not build it)
DDI validation assets — DDI 2.5 XSDs + the codegen-written ddi_custom_rules.sch ddi-validation/{xsd,schematron}/ qwacback, synced via its .registry-version pin
cdl-survey-types skill — self-contained Agent Skill (SKILL.md + references/), generated from the registry; portable to any agent runtime that reads the format skills/cdl-survey-types/ formulaid's generating-xlsforms skill, which owns the workflow and includes this as the type reference

Everything else here (docs/ spec site, screenshots/, tests, fixtures) is in-repo only.

Installation

As a library

Each release carries a prebuilt package. Installing it runs no build step:

npm install https://github.com/CorrelAid/formtransform/releases/download/v0.1.7/correlaid-formtransform-0.1.7.tgz

Installing from git also works, but builds dist/ on install through the prepare script, which needs TypeScript and install scripts enabled:

npm install github:CorrelAid/formtransform#v0.1.7

Releases also attach cdl-survey-types-<version>.tar.gz, the generated skills/cdl-survey-types/ sub-skill, and formtransform-fixtures-<version>.tar.gz, the example fixtures and blessed snapshots (registry/entities/, tests/fixtures/surveys/, same paths) for golden tests. Assets never change after publishing, so their checksums can be pinned: each release also attaches SHA256SUMS and lists the checksums in its notes.

As a CLI tool

npx github:CorrelAid/formtransform --help

Quick Start

TypeScript/JavaScript

One function per direction. An XLSForm source is the .xlsx bytes or its loaded sheets ({ surveyData, choicesData, settingsData }); a LimeSurvey source is the structure TSV text:

import {
  xlsformToLstsv,
  xlsformToDdi,
  lstsvToDdi,
  lstsvToXlsform,
  ddiToXlsform,
} from '@correlaid/formtransform';

const bytes = await file.arrayBuffer();
const tsv = await xlsformToLstsv(bytes);
const xml = xlsformToDdi(bytes);
const fromLs = lstsvToDdi(tsv);
const { survey, choices, settings } = lstsvToXlsform(tsv);
const back = ddiToXlsform(xml); // a codebook, or a <var>/<varGrp> fragment

Each rejects input outside its subset with a ConversionError unless you pass skipValidation: true (see Supported XLSForm Subset).

The building blocks behind these (the converter class, type mapper, TSV serializer, expression transpiler) are exported from @correlaid/formtransform/internals, with no stability promise. XLSFormToTSVConverter, XLSFormParser, buildDdiXml and lstsvToDdiXml still work from the main entry point but are deprecated and will be removed from it in the next minor release.

A DDI codebook and the data file it describes come from the same variable list, so the CSV headers match the XML <var name=""> elements one-to-one. For the whole path from a Kobo or LimeSurvey export, see RESPONSE_DATA.md:

import {
  xlsformToDdi,
  buildDataCsv,
  extractVariables,
  choicesByListFromRows,
} from '@correlaid/formtransform';

const xml = xlsformToDdi({ surveyData: survey, choicesData: choices, settingsData: settings }, { submissions });
const csv = buildDataCsv(
  extractVariables(survey, choicesByListFromRows(choices)),
  submissions,
);

Command Line

# Convert XLSForm to LimeSurvey TSV
formtransform xlsform2lstsv survey.xlsx -o survey.tsv

# Convert to DDI Codebook
formtransform xlsform2ddi survey.xlsx -o codebook.xml

# ...plus the response-data CSV (flat CSV, `;` or `,`, or a Kobo submissions
# JSON array). Writes data.csv beside codebook.xml and sets <caseQnty>.
formtransform xlsform2ddi survey.xlsx -o codebook.xml --data responses.csv

# Same for LimeSurvey: structure TSV + response export (question-code headings)
formtransform lstsv2ddi survey.tsv -o codebook.xml --data responses.csv

# Convert from LimeSurvey TSV back to XLSForm (emitted as JSON sheets)
formtransform lstsv2xlsform survey.tsv -o recovered.json

# Convert a DDI codebook (or a <var>/<varGrp> fragment) back to XLSForm (JSON sheets)
formtransform ddi2xlsform codebook.xml -o form.json

Asking the library what exists

Two generated catalogues are part of the public API, for consumers that render or generate surveys rather than convert them:

import { QUESTION_TYPES, APPEARANCES, TYPE_MAPPINGS } from '@correlaid/formtransform';

QUESTION_TYPES.select_one.label;          // "Select One"
QUESTION_TYPES.select_one.labels.de;      // "Einfachauswahl"
QUESTION_TYPES.select_one.useWhen;        // when to reach for this type
QUESTION_TYPES.select_one_other.base;     // "select_one" — a variant of it
QUESTION_TYPES.select_one_other.presentation; // { withOther: true, withLongList: false }
QUESTION_TYPES.grid.bases;                // composites span several types
QUESTION_TYPES.select_one.constraints;    // name/choice-code limits
APPEARANCES.label.carriesData;            // false — a matrix header stores no answer
TYPE_MAPPINGS.select_one.limeSurveyType;  // "L" — how it converts

QUESTION_TYPES is keyed by registry slug and answers what a row can be (labels in English and German, guidance, variant → base and presentation, authoring constraints, metadata rows); TYPE_MAPPINGS answers how it converts. Both are generated from registry/ — never hardcode a type list or a label in a consumer.

The Registry is the Source of Truth

registry/ is the single source every other artifact in this repo derives from. It is hand-authored (JSON-LD + vocabulary CSVs + per-entity fixtures); nothing else here defines what a question type is, how it maps between standards, or what is allowed. Every downstream artifact — the TypeScript library and its CLI, the generated Schematron rules and the worker image that ships them, the Claude Code skill, the spec site, the test fixtures — is either generated from the registry by codegen or reads registry-derived data at runtime.

Concretely, this means:

  • A behaviour change starts in registry/, never downstream. Editing generated files by hand is pointless — the next uv run codegen overwrites them.
  • codegen is the only writer of those artifacts, and it validates the registry first, so an invalid registry cannot produce artifacts at all.
  • Downstream repos inherit the registry transitively. formtransform-app and qwacback depend on @correlaid/formtransform and on the worker image; formulaid depends on the generated skill. None of them carry their own type list — adding a question type here is what makes it exist for all of them.

Supported Standards

Three standards, each owned by a different tool and built for a different job. They disagree on basic terms — what a question is, what its parts are — so the library routes every transformation through the canonical registry, which records how each maps onto it.

Standard Role Type system
XLSForm (Kobo/ODK) Authoring — surveys are written here type strings (select_one, integer)
LimeSurvey TSV Deployment — recreate the survey in LimeSurvey type codes (L, M, F, N)
DDI Codebook 2.5 Documentation — describe the resulting dataset interval class + response domains (category, multiple)

Supported directions: five, one per module under src/pipelines/ — xlsform2lstsv (deploy the survey), xlsform2ddi (document the dataset), lstsv2ddi, lstsv2xlsform and ddi2xlsform (the reverse paths). All are lossy for some types: nested groups flatten in LimeSurvey, choice codes over 5 chars truncate, select_multiple becomes N binary variables, and the LimeSurvey reverse paths cannot recover a select's authored list_name. A CDL codebook holds the whole form: XLSForm → DDI → XLSForm gives it back, and its codebook again.

DDI goes back to XLSForm: a CDL codebook carries the whole form. Skip logic, validation and required are each a readable <universe> sentence (a simple numeric range also <valrng>), plus the exact expression in a typed note such as <notes type="cdl:relevant" subject="xlsform-xpath"> (convention:logicMapping). Groups, order, hints, defaults, appearances and parameters, list names, note rows and metadata rows, settings and language names are in standard DDI where it has a place and typed notes where not (convention:ddiFields). ddiToXlsform reads it back; any other DDI converts as far as its standard elements go, with a warning per missing field (src/pipelines/ddi2xlsform/README.md).

Errors and warnings

Every error the library throws is a ConversionError with a stable code (type-unregistered, choice-list-empty, xpath-syntax, …; see DiagnosticCode), a human message, and subject, the question it concerns. Branch on code, not on the message text.

xlsformToLstsv and xlsformToDdi check the whole form against the subset first (validateSubset, see below). If anything is outside it, they throw one ConversionError with code xlsform-outside-subset, and details holds every finding, so a form with several problems is fixed in one pass.

Warnings (a truncated name, an ignored appearance, a comparison with a value that isn't one of the choices, a constraint that can't be converted) go to an onWarning callback. This includes the subset check's warnings, each passed once. The default prints them to the console:

const warnings: Diagnostic[] = [];
await xlsformToLstsv(bytes, { onWarning: (w) => warnings.push(w) });

validateSubset returns the same Diagnostic objects without converting.

Supported XLSForm Subset

Not everything XLSForm allows is registered (supported). The library strictly validates against this subset, rejecting:

  • Unregistered types (no LimeSurvey equivalent — image, audio, geopoint, etc.)
  • Unregistered appearances (warn + ignore)
  • Out-of-subset names — field names and answer codes must match ^[a-zA-Z0-9]+$ (no underscores, except the <question>_other companion), names at most 20 characters, codes at most 5. FieldSanitizer turns free text into conforming names: it transliterates (ä→ae, ß→ss), drops other diacritics and deletes the rest of the non-alphanumerics
  • Deep nesting — max 3 levels (group/group/question)
  • Duplicate or missing answer codes — each choice needs a code, unique within its list (an empty choice label is only a warning)
  • Unresolvable answer options — a select_one/select_multiple needs a list name with rows on the choices sheet; a select_*_from_file needs a registered vocabulary (e.g. iso_3166_1.csv) or a CSV passed as fileChoices (the CLI reads CSVs beside the form)
  • Unmappable languages — language tags are BCP 47 (de, fr-BE, zh-Hans), in label::<tag> or label::Name (<tag>) columns and default_language. LimeSurvey has a fixed code list: a regional tag it lacks becomes its language with a warning (fr-BE → fr), and a language it lacks entirely (eo), or two tags on one code, is an error. DDI keeps tags as written and carries every language the form has: the base language (default_language, else the first) untagged and first, each other one as an xml:lang sibling. Only the form's own texts go in, never a translation; a language that lacks a text gets no element
  • Reserved words — relevance, validation, text, etc. (LimeSurvey internals)
  • Dangling references — every ${name} in relevant or constraint must name a row of the survey sheet. A compared literal that the question can never take (${q} = 'Sonstiges' when the code is sonst, selected(${m}, 'x'), ${age} = 'old') is only a warning: the form converts, but the condition is never true

An exclusive answer in a select_multiple ("Keine Angabe", "Nichts davon") is marked with an exclusive column (yes) on the choices sheet, not a count-selected() constraint. LimeSurvey enforces it through exclude_all_others; DDI has no field for it.

The or_other shorthand writes no text for its "other" answer; the survey tool supplies one. LimeSurvey shows its own Other: in the survey language (Sonstiges:), ODK and Kobo an untranslated "Other" / "Specify other.". The DDI records LimeSurvey's wording in each of the form's languages, and the DDI check warns (other-shorthand). To record your own text on every platform, write an other choice and a <question>_other text question instead.

The name and code limits are LimeSurvey's. For DDI, check with validateSubset(survey, choices, { target: 'ddi' }) (CLI: validate --target ddi; xlsform2ddi does it by default): same rules without those limits, since DDI keeps names as authored. xlsformToDdi runs this check itself (skipValidation turns it off).

LimeSurvey's reverse-subset check (lstsv2xlsform) is narrower — no arrays, no ranking, no numeric/date expressions.

Development

# Install dependencies
npm install
uv sync

# Run tests
npm test           # TypeScript tests
uv run pytest      # Python tests
npm run test:live  # Docker integration tests

# Generate artifacts from registry
uv run codegen

# Bless snapshots after registry changes
npm run bless

Releasing

  1. Bump version in package.json in a PR and merge it.
  2. Publish a GitHub release tagged v<version> on that commit.

Publishing the tag builds the schematron-worker image (worker-image.yml). Publishing the release attaches the package tarball and the skill archive (release-assets.yml). That job fails if the tag and package.json disagree, and it never replaces an asset that already exists: consumers pin the checksums, so a fix ships as a new version.

Documentation

  • Architecture — Technical architecture and internal structure
  • Response data — Kobo/LimeSurvey export → DDI codebook + data CSV (CLI and browser), and reading the codebook back in Python
  • survey2ddi handover — plan for retiring survey2ddi in favour of this library (lives in survey2ddi)
  • Question-bank alignment: tracked in qwac (#9, #10, #12, #13) and qwacback (#3 converter, #4 validation worker)
  • Website content (umfragen.civic-data.de) that advertises these tools and feeds agent-readable XLSForm docs: tracked in cdl-wp-eins #26–#31
  • Claude Code Integration — Claude-specific features and skills
  • Pipeline Documentation — Transformation pipeline details
  • Test Documentation — Test structure and running tests

AI Disclosure

See AI_DISCLOSURE.md.

License

MIT

About

Registry-driven survey transformation between XLSForm, LimeSurvey TSV and DDI Codebook 2.5

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages