Population-genetics-grounded genomic selection dataset: 400 genotypes, 800 SNP-style markers with realistic linkage disequilibrium and pedigree relatedness, 12 environments with developmental-stage-resolved covariates, 4 traits with additive + dominance + epistasis + pleiotropy + reaction-norm G×E architecture, and ground-truth genetic effects for benchmarking prediction accuracy against known truth.
CropGxE-GenomicSelection-Synthetic simulates a 400-line breeding population genotyped at 800 SNP-style markers across 10 chromosomes, generated with realistic linkage disequilibrium (a haplotype-copying process, not independent random markers) and real pedigree relatedness (simulated biparental crosses with actual meiotic recombination). Four traits (grain yield, days to flowering, disease resistance score, plant height) follow a major-QTL-plus-polygenic-background genetic architecture, extended with dominance deviations, additive-by-additive epistasis, and deliberately shared (pleiotropic) QTL between trait pairs to induce genetic correlation. A trait-specific fraction of QTL are environment-plastic, modulated by developmental-stage-appropriate environmental covariates -- flowering time responds to temperature at flowering specifically, yield responds to rainfall during grain-fill, not season-long averages. The genomic relationship matrix is computed with the standard VanRaden method used throughout real genomic prediction. Multi-environment trials use a realistic sparse design supporting the standard cross-validation schemes used in genomic selection research. Unlike any real breeding trial dataset, this dataset ships the ground-truth QTL, dominance, and epistatic effects used to generate every phenotype, enabling direct benchmarking of a prediction method's estimates against known truth.
This Dataset Supports Research On Genomic Selection And Genotype-by-environment Interaction In Crop Breeding, An Area That In India Currently Exists Mainly As Closed Institutional Data Rather Than An Open, Structurally Realistic Benchmark. It Targets Researchers Comparing Additive-only Genomic Prediction Models (Standard Gblup) Against Models That Also Estimate Dominance, Epistasis, And Genotype-by-environment Effects, Since The Ground-truth Architecture Lets A Method's Accuracy Be Checked Against Known Truth Rather Than Only Cross-validated Correlation With Noisy Phenotypes. It Also Supports Research On Multi-trait Genomic Prediction, Since Traits Are Not Generated Independently But Carry Deliberate Genetic Correlations, And On Cross-environment Generalization Using The Dataset's Realistic Sparse Multi-environment Trial Design.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.