Model the schema and workflow first
My wellbeing dissertation created persistent student IDs, academic terms, WHO-5 records, engagement metrics and safeguarding flags so the prototype could exercise realistic joins and longitudinal views.
Encode constraints deliberately
Synthetic data is more useful when values obey domain rules, such as valid score ranges, non-negative usage metrics and consistent temporal relationships.
Do not confuse plausibility with validation
Synthetic labels can test a pipeline but cannot prove predictive validity on real populations. Performance metrics against simulated targets are engineering evidence, not clinical or operational validation.
Use it as a roadmap
A strong synthetic dataset can reveal which telemetry, identifiers and timestamps a future production system needs to collect.