Editing
Portal:Life Sciences
(section)
Jump to navigation
Jump to search
Warning:
You are not logged in. Your IP address will be publicly visible if you make any edits. If you
log in
or
create an account
, your edits will be attributed to your username, along with other benefits.
Anti-spam check. Do
not
fill this in!
=== The Human Genome Project era === The Human Genome Project, carried out from 1990 to 2003, was one of the most important scientific projects in modern biology. Its goal was to generate the first sequence of the human genome, along with sequences of several important model organisms. The National Human Genome Research Institute describes the Human Genome Project as a large, organized, international effort that generated the first human genome sequence and helped establish open data-sharing practices in genomics.<ref>National Human Genome Research Institute, “Human Genome Project Fact Sheet,” https://www.genome.gov/about-genomics/educational-resources/fact-sheets/human-genome-project, accessed July 3, 2026.</ref> Perl was heavily used during this period because genome sequencing was not only a laboratory problem. It was a data-management problem. Sequencing centers generated enormous numbers of reads, quality values, clone names, map positions, assembly files, and annotations. These data arrived in many formats and changed as techniques changed. At the Washington University Genome Sequencing Center, early large-scale sequence preprocessing work used Unix-based Perl systems. A 1998 Genome Research paper on automated sequence preprocessing stated that the group had begun developing its strategy and Unix-based Perl software for large-scale sequence preprocessing at the Genome Sequencing Center.<ref>Michael C. Wendl et al., “Automated Sequence Preprocessing in a Large-Scale Sequencing Environment,” Genome Research, 1998, https://pmc.ncbi.nlm.nih.gov/articles/PMC310779/, accessed July 3, 2026.</ref> This kind of work was crucial. Before a DNA sequence could become part of a reference genome or a biological conclusion, raw data had to be cleaned, checked, named, assembled, tracked, and connected to other data. Perl helped automate those steps. An often-cited BioPerl article, “How Perl saved the Human Genome Project,” argued that Perl filled the needs of genome centers by handling incompatible data formats, rapidly changing techniques, and monolithic analysis programs. It described Perl as a practical tool that genome centers frequently turned to when they needed to solve problems quickly.<ref>BioPerl, “How Perl saved the Human Genome Project,” https://bioperl.org/articles/How_Perl_saved_human_genome.html, accessed July 3, 2026.</ref> The title is intentionally dramatic, but the underlying point is historically fair: Perl was one of the main scripting languages that made early genome-center data processing manageable.
Summary:
Please note that all contributions to Perl Guilds - Getting Medieval with Perl may be edited, altered, or removed by other contributors. If you do not want your writing to be edited mercilessly, then do not submit it here.
You are also promising us that you wrote this yourself, or copied it from a public domain or similar free resource (see
Perl Guilds - Getting Medieval with Perl:Copyrights
for details).
Do not submit copyrighted work without permission!
Cancel
Editing help
(opens in new window)
Navigation menu
Personal tools
Not logged in
Talk
Contributions
Log in
Namespaces
Page
Discussion
English
Views
Read
Edit
Edit source
View history
More
Search
Navigation
Main page
Recent changes
Random page
Help about MediaWiki
Special pages
Tools
What links here
Related changes
Page information