Genomics teams run sample sheets, analysis, storage, consent records, partner deliverables and field trials in separate spreadsheets and systems, so one index clash or one withdrawn consent spreads before anyone can see how far.
A modern biotech runs on throughput. Libraries are pooled onto flow cells by the hundred, analysis pipelines turn raw reads into variant calls and expression tables, and the data grows faster than any budget line. Much of that flow is still stitched together by hand: sample sheets built in spreadsheets, run quality checked in instrument folders, pipeline versions remembered rather than recorded, and storage spread across an on-premises cluster and several cloud accounts, with no one sure what each terabyte belongs to.
Around the data sit obligations that do not forgive gaps. Human genomic data carries the consent each participant gave, and that consent differs by form version, by use and by partner. Collaborations with pharmaceutical and seed partners pay on milestones and fund an agreed number of scientists, and partners expect evidence rather than slides. Agricultural trait programs run regulated field trials whose permits require isolation from compatible crops, a controlled harvest and every seed packet accounted for.
The pressure lands on a handful of people: the core manager who learns of a failed lane when submitters start emailing, the bioinformatics lead asked to re-run a thousand samples on a new reference build, the data steward handling a withdrawal in the same week as two access requests, the alliance manager preparing a steering committee, and the trials manager racing rain to a harvest window. Each is solving a problem that crosses teams with a tool built for one.