Ninety percent of the work happens in a database nobody outside the trial ever looks at. That is the honest starting point. Sponsors talk about endpoints and p-values, but before any of that means anything, someone has to turn a stack of case report forms into a dataset a regulator will accept. I have sat through enough query resolution cycles to know the difference between a trial that finishes on schedule and one that drags is almost never the science. It is the housekeeping.
So here is what actually happens between a patient signing consent and a statistical analysis plan producing a table. Most of it is unglamorous, and all of it matters more than people outside the industry realize.
If you are running a study and you have not thought hard about who handles your data pipeline, that gap is where timelines go to die. Teams that bring in Clinical Data Management and Biostatistics support early tend to catch structural problems before they harden into months of rework, because the people cleaning the data and the people analyzing it are reading from the same playbook.
What data management actually covers
People hear “data management” and picture a spreadsheet. It is closer to air traffic control. The job spans database design, edit check programming, discrepancy tracking, medical coding, and the locked final dataset that gets handed off for analysis. Every one of those steps produces a deliverable someone downstream depends on.
Database design comes first, and it is where you win or lose. If your electronic data capture forms do not map cleanly to the protocol’s endpoints, you will spend the rest of the trial compensating. I have watched teams rebuild a mid-study database because a scale was entered as a free text field instead of a controlled number. That is a two month mistake caused by a thirty minute oversight.
Edit checks come next. These are the rules that flag impossible or suspicious entries: a blood pressure reading of 400 over 200, a visit date that predates consent, a lab value outside the analyzable range. Good edit checks catch errors at the source. Bad ones drown your monitors in false positives until everyone starts ignoring the list, which is worse than having no list at all.
Query resolution is where your timeline lives
A query is a question sent back to the site about a data point that does not add up. The cycle looks simple. The coordinator reviews it, answers it, and the data manager closes it. In practice, queries bounce, sites get busy, and a single unresolved item can hold an entire patient record hostage.
Here is the part sponsors underestimate. Query volume is not a quality problem. It is a design problem. A trial generating three queries per patient usually has ambiguous form fields, not sloppy coordinators. I would rather see fifty well targeted queries than five vague ones, because targeted queries close fast and teach the site something for next time.
My rule of thumb after running these cycles: track age of open queries weekly, not total count. Total count tells you how busy you are. Query age tells you whether you will actually lock the database when you promised.
Why biostatistics enters earlier than people think
There is a persistent myth that statisticians show up at the end, run the numbers, and hand over tables. That workflow produces clean math on messy questions.
Statisticians belong in the room when the protocol is drafted. They decide the sample size, define the analysis populations, and specify how missing data gets handled. Those three decisions shape everything. If the primary endpoint analysis is not nailed down before the first patient enrolls, you are guessing at the finish line while already running the race.
The regulatory expectation here is not a secret. The U.S. Food and Drug Administration expects statistical methods to be pre-specified, and reviewing its guidance on trial design is the fastest way to see how much scrutiny those choices receive.
Practical translation: freeze your statistical analysis plan before unblinding. Changing it after you see the data invites a credibility problem you cannot argue your way out of.
A concrete scenario worth stealing from
A mid-sized sponsor I worked alongside was running a metabolic trial across three sites. Enrollment finished on time. Then the database refused to lock for eleven weeks. The culprit was not one dramatic failure. It was four categories of small ones:
- Visit windows calculated differently at two sites because the form never defined which date counted as day one
- Concomitant medications logged with brand names at one site and generics at another, slowing medical coding
- Missing lab values recorded as “not done” in a field that only accepted numbers
- One site answering queries two weeks slower than the others, invisible until someone finally charted response times by location
None of these were scientific problems. All of them were preventable at the design stage. The fix that eventually worked was a shared data conventions document, built before the first site activation, that spelled out date logic, coding standards, and who owns each field. The next trial used it and locked the database within three weeks of last patient visit.
Standards exist so you do not reinvent the wheel
You do not have to invent your own conventions from scratch. Industry bodies have spent years building the frameworks, and ignoring them means paying for lessons your competitors already learned.
The International Council for Harmonisation publishes the Good Clinical Practice guidelines that shape how data integrity is judged worldwide. Aligning your documentation with those expectations is not bureaucracy. It is the difference between a clean inspection and a long one.
Beyond the standards, the skill that separates strong data teams is traceability. Can you show a regulator exactly how a number moved from a source document to a final table, with every transformation documented? If the answer takes more than a few minutes to find, your process has a hole in it. The National Institutes of Health treats data sharing and documentation rigor as baseline expectations for the research it supports, which tells you where the bar sits.
What to do before your next trial starts
You do not need a bigger team. You need decisions made in the right order.
- Draft the data conventions document before site activation, not after the first query storm.
- Have your statistician sign off on endpoints, populations, and missing data rules while the protocol is still editable.
- Build edit checks that a coordinator can act on without calling anyone.
- Track query age weekly and report it to leadership with the same seriousness as enrollment numbers.
- Rehearse your database lock as a scheduled milestone, with a named owner and a real date.
Pick the one that scares you most and start there. For most teams, that is number two, because it forces a conversation with the statistician earlier than feels comfortable. Do it anyway.
Clean data is not the goal of a trial. It is the price of admission for anyone who wants their results to be believed. Set that standard at the design table, and the rest of the trial gets a lot quieter. Where does your current process break first, the form design or the query cycle?

