{"title": "Illumina (short-read)", "status": "open", "content": [{"body": "## Overview\n\nThe paired-end short-read alignment pipeline for Illumina data follows the Genome Analysis Toolkit (GATK) Best Practices. It is designed for per-sample and per-library execution, handling one or multiple sets of paired FASTQ files. The pipeline is optimized for distributed processing, requiring each pair of FASTQ files to correspond to a single sequencing lane.\n\n### Key Pipeline Steps\n\n1. **Alignment with BWA-MEM:** Initial alignment of the raw reads to the reference genome using BWA-MEM.\n2. **Read Groups Assignment:** Assignment of reads to specific groups, providing essential information for subsequent steps.\n3. **Duplicate Reads Marking:** Identification and labeling of duplicate reads, that originated during library preparation or as sequencing artifacts.\n4. **Local Indel Realignment:** Local realignment at identified indel positions.\n5. **Base Quality Score Recalibration (BQSR):** Recalibration of base quality scores, including indels, resulting in the generation of a refined, analysis-ready BAM file.\n\nA modified alignment step with BWA-MEM is used for processing Hi-C data.\n\n### Sentieon Software and Distributed Mode\n\nTo meet the scalability demands, especially with high-coverage whole-genome sequencing data, the pipeline utilizes a more efficient software implementation from [Sentieon](https://www.sentieon.com/).\n\nSentieon offers a comprehensive toolkit (DNASeq<sup><sub>1</sub></sup>) that replicates the original BWA and GATK algorithms, while enhancing computational efficiency. The pipeline also leverage Sentieon\u2019s [distributed mode](https://support.sentieon.com/appnotes/distributed_mode/) to streamline the processing of large data volumes.\n\nThe distributed version of the pipeline operates on reads divided by sequencing lanes and parallelizing some of the processing across multiple genomic regions, defined as shards. This implementation is a 1 to 1 equivalent to a standard implementation of the pipeline, producing identical results.\n\n### Pipeline Chart\n\n![flow_chart](/static/img/pipeline-docs/Flow_Chart_Pipeline_Short-Read_Illumina_Paired-End.png)\n\n<sub><b>1</b>: *Kendig KI, Baheti S, Bockol MA, Drucker TM, Hart SN, Heldenbrand JR, Hernaez M, Hudson ME, Kalmbach MT, Klee EW, Mattson NR, Ross CA, Taschuk M, Wieben ED, Wiepert M, Wildman DE, Mainzer LS.* Sentieon DNASeq Variant Calling Workflow Demonstrates Strong Computational Performance and Accuracy. *Front Genet. 2019 Aug 20;10:736.* d\noi: 10.3389/fgene.2019.00736</sub>\n\n---\n\n## Alignment with BWA-MEM\n\nFor the initial alignment, the pipeline uses BWA-MEM in paired-end mode on each set of paired FASTQ files. The reads are then sorted by genomic coordinates, and an integrity check is performed on the resulting BAM file.\n\n### Aligning and Sorting\n\n###### Align and sort reads\n\n<pre class=\"code-block copy-wrapper\">\nsentieon bwa mem -K 10000000 reference.fasta reads.fastq mates.fastq |\n  samtools sort -o sorted.bam -\n</pre>\n\nArguments:\n\n- *-K*: chunk size option to have number of threads independent results.\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nSentieon implementation replicates the original [BWA](https://github.com/lh3/bwa) code. The pipeline is using Sentieon version 202308.01, corresponding to BWA version 0.7.17.\n\n*Note: In the original BWA code there is a programming error affecting alignment score calculations. Specifically, when secondary alignments are present, the score calculation includes unrelated information, potentially influencing the Mapping Quality (MAPQ) of the primary alignment by lowering it. Sentieon introduced a fix to the issue, resulting in slight differences in MAPQ between the two software (approximately 0.008% of reads are affected).*\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_bwa-mem_sort.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_bwa-mem_sort.sh) [BWA-MEM]\n\n---\n\n## Read Groups\n\nAs the second step, the pipeline assigns the reads to unique read groups, representing identifiers that group reads together. A read group (`@RG`) captures relevant information about the sample and the sequencing process and technology, utilized by various downstream bioinformatics tools.\n\nThe relevant fields in defining a read group include:\n\n- **ID (Identifier):** A unique identifier for the read group within the BAM file and across multiple BAM files used in the same dataset.\n- **SM (Sample):** The sample to which the reads belong.\n- **PL (Platform):** The technology used to sequence the reads. This tag is required if running Base Quality Score Recalibration (BQSR) to determine the correct error model (e.g., ILLUMINA).\n- **PU (Platform Unit):** A unique identifier for the sequencer unit used for sequencing (i.e., sequencing lane). This tag is required if running BQSR, as it models together all reads belonging to the same platform unit.\n- **LB (Library):** The library used to sequence the reads. This tag is used in the process of marking or removing duplicate reads to determine groups that may contain duplicates, as duplicate reads need to belong to the same library.\n\n### Assigning Read Groups\n\nTo assign read groups, an in-house Python script is used. It can automatically generate read groups based on Illumina read names and handle multiple read groups in the same file (e.g., reads from multiple lanes are merged into a single file).\n\nThe read groups are assigned as follows:\n\n- **ID:** `<sample name>.<instrument>_<run>_<flow cell>.<lane>`\n- **SM:** `<sample name>`\n- **PL:** `<platform>`\n- **PU:** `<instrument>_<run>_<flow cell>.<lane>`\n- **LB:** `<sample name>.<library>`\n\nE.g., in BAM file:\n\n<pre class=\"code-block copy-wrapper\">\n@RG ID:SMAHT1.ST-E00127_336_HJ7YHCCXX.8  SM:SMAHT1  PL:ILLUMINA  PU:ST-E00127_336_HJ7YHCCXX.8  LB:SMAHT1.HISEQ-LIB1\n</pre>\n\n### Source Code\n\nAll the relevant code is accessible in the GitHub repository:\n\n  - [AddReadGroups.py](https://github.com/smaht-dac/pipelines-scripts/blob/main/processing_scripts/AddReadGroups.py) [AddReadGroups]\n\n---\n\n## Duplicate Reads\n\nIn this step, the pipeline marks duplicate reads. Duplicate reads are sequencing artifacts that originate during library preparation and sequencing runs. Duplicate reads are evaluated per-library using the `LB` tag in the read groups.\n\nThe pipeline does not remove the duplicate reads that are tagged directly in the BAM file.\n\n### Detecting and Marking Duplicates\n\n###### Detect duplicate reads\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -i sorted.bam\n                --algo LocusCollector\n                --fun score_info\n                score.txt\n</pre>\n\n###### Mark duplicate reads\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -i sorted.bam\n                --algo Dedup\n                --optical_dup_pix_dist 2500\n                --score_info score.txt\n                deduped.bam\n</pre>\n\nArguments:\n\n- *-\\-optical_dup_pix_dist*: maximum offset between two duplicate clusters to consider them optical duplicates. For structured flow cells (NovaSeq, HiSeq 4000, X), the pipeline uses 2500.\n\n*Note: The actual implementation of the above command in the pipeline is more complex to support distributed execution, but functionally equivalent.*\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nThe pipeline implementation uses Sentieon LocusCollector to calculate duplicate metrics per library and the Dedup algorithm to mark duplicate reads in the BAM file. The pipeline is using Sentieon version 202308.01, corresponding to Picard 2.9.0. Both algorithms combined are equivalent to the MarkDuplicates algorithm in Picard.\n\n###### Detect and mark duplicate reads (Picard equivalent)\n\n<pre class=\"code-block copy-wrapper\">\njava -jar picard.jar MarkDuplicates\n      INPUT=sorted.bam\n      OPTICAL_DUPLICATE_PIXEL_DISTANCE=2500\n      OUTPUT=deduped.bam\n</pre>\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_LocusCollector.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_LocusCollector.sh) [LocusCollector]\n  - [sentieon_LocusCollector_apply.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_LocusCollector_apply.sh) [Dedup]\n\n---\n\n## Local Realignment\n\nIn this step, the pipeline perform local realignment at indel positions. Unlike genome aligners that consider each read independently and may favor alignments with mismatches or soft-clips over opening a gap, local realignment takes into account all reads spanning a given position, enabling a high-scoring consensus that supports the presence of an indel event. The two-step realignment process identifies potential regions for alignment improvement and then realigns the reads in these regions using a consensus model that considers all reads in the alignment context together. This allows to correct potential mapping errors made by the aligners, enhancing the consistency of read alignments in regions with indels.\n\n### Realigning Reads\n\n###### Realign reads around indels\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -r reference.fasta\n                -i deduped.bam\n                --algo Realigner\n                -k known_sites_INDEL.vcf\n                realigned.bam\n</pre>\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nThe implementation uses Sentieon Realigner algorithm to identify the potential targets and perform local indel realignment. The pipeline is using Sentieon version 202308.01, corresponding to GATK versions 3.7, 3.8. The algorithm is equivalent to RealignerTargetCreator and IndelRealigner algorithms in GATK.\n\n###### Identify potential targets (GATK equivalent)\n\n<pre class=\"code-block copy-wrapper\">\njava -jar GenomeAnalysisTK.jar \n            -T RealignerTargetCreator\n            -R reference.fasta\n            -I deduped.bam \n            -known known_sites_INDEL.vcf\n            -o realignment_targets.list\n</pre>\n\n###### Local indel realignment (GATK equivalent)\n\n<pre class=\"code-block copy-wrapper\">\njava -jar GenomeAnalysisTK.jar \n            -T IndelRealigner\n            -R reference.fasta\n            -I deduped.bam\n            -targetIntervals realignment_targets.list\n            -known known_sites_INDEL.vcf\n            -o realigned.bam\n</pre>\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_Realigner.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_Realigner.sh) [Realigner]\n\n---\n\n## Base Quality Score Recalibration (BQSR)\n\nIn this final step, the pipeline recalibrates the base quality scores produced by the sequencing machine, correcting systematic (non-random) technical errors leading to over- or under-estimated base quality scores in the data. BQSR modeling is independently performed on groups of reads from different sequencer units (i.e., lanes), identified by the `PU` tag in the read groups.\n\n### Base Quality Score Modeling and Recalibration\n\n###### Model quality scores\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -r reference.fasta\n                -i realigned.bam\n                --algo QualCal\n                -k known_sites_SNP.vcf\n                -k known_sites_INDEL.vcf\n                recal_data.table\n</pre>\n\n###### Apply score recalibration\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -r reference.fasta\n                -i realigned.bam\n                --read_filter 'QualCalFilter,table=recal_data.table,keep_oq=true'\n                --algo ReadWriter\n                recalibrated.bam\n</pre>\n*Note: keep_oq=true in the -\\-read_filter argument will preserve the original base quality scores by using the OQ tag in the recalibrated bam file.*\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nThe implementation uses Sentieon QualCal algorithm to construct models of covariation based on the input data and a set of known variants. This process produces the recalibration table necessary for BQSR. The recalibration is then applied to the BAM file using the Sentieon ReadWriter command. The pipeline is using Sentieon version 202308.01, corresponding to GATK versions 3.7, 3.8, 4.0, and 4.1. The algorithms are equivalent to BaseRecalibrator and ApplyBQSR algorithms in GATK.\n\n###### Model quality scores (GATK equivalent)\n\n<pre class=\"code-block copy-wrapper\">\ngatk BaseRecalibrator -R reference.fasta\n                      -I realigned.bam\n                      --enable-baq\n                      --known-sites known_sites_SNP.vcf\n                      --known-sites known_sites_INDEL.vcf\n                      -O recal_data.table\n</pre>\n\n###### Apply score recalibration (GATK equivalent)\n\n<pre class=\"code-block copy-wrapper\">\ngatk ApplyBQSR -R reference.fasta\n               -I realigned.bam\n               -bqsr recal_data.table\n               -O recalibrated.bam\n</pre>\n\n*Note: To reduce run time (at the expense of accuracy), GATK4 disables the base quality score recalibration of indels when using default settings. Consequently, the Sentieon BAM output will contain BI/BD tags from the indel recalibration, which will be missing from the GATK4 BAM output when run with default settings.*\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_QualCal.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_QualCal.sh) [QualCal + ReadWriter]\n\n---\n\n## Hi-C\n\nFor the alignment of Hi-C data, the pipeline uses BWA-MEM in paired-end mode on each set of paired FASTQ files. The reads are then sorted by genomic coordinates, and an integrity check is performed on the resulting BAM file. Compared to standard alignment, extra flags are required to support the read pairs generated for the Hi-C data.\n\n### Aligning and Sorting\n\n###### Align and sort reads\n\n<pre class=\"code-block copy-wrapper\">\nsentieon bwa mem -5SP -K 10000000 reference.fasta reads.fastq mates.fastq |\n  samtools sort -o sorted.bam -\n</pre>\n\nStandard arguments:\n\n- *-K*: chunk size option to have number of threads independent results.\n\nHi-C specific arguments:\n\n- *-S*: skip mate rescue.\n- *-P*: skip pairing. Mate rescue performed unless -S also in use.\n- *-5*: for split alignment, take the alignment with the smallest coordinate as primary.\n\n*-SP* ensures that the results are equivalent to aligning each mate separately, while maintaining the proper paired-end read formatting. It also avoids forced alignment of a poorly aligned read given an alignment of its mate, based on assumption that the two mates belong to a single genomic segment.\n\n*-5* ensures that the 5' portion of chimeric alignments is marked as the primary alignment. In Hi-C experiments, the 5' portion of the read is typically the alignment of interest, with the 3' portion representing the same fragment as the mate.\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nSentieon implementation replicates the original [BWA](https://github.com/lh3/bwa) code. The pipeline is using Sentieon version 202308.01, corresponding to BWA version 0.7.17.\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_bwa-mem_sort_Hi-C.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_bwa-mem_sort_Hi-C.sh) [BWA-MEM]", "title": "Illumina (short-read)", "status": "open", "options": {"filetype": "md", "collapsible": false, "default_open": true, "convert_ext_links": true, "initial_header_level": 2}, "consortia": [{"display_title": "SMaHT", "@type": ["Consortium", "Item"], "uuid": "358aed10-9b9d-4e26-ab84-4bd162da182b", "status": "open", "@id": "/consortia/358aed10-9b9d-4e26-ab84-4bd162da182b/", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "identifier": "short-read_illumina_paired-end", "date_created": "2026-01-09T20:07:18.862130+00:00", "section_type": "Page Section", "submitted_by": {"error": "no view permissions"}, "last_modified": {"modified_by": {"error": "no view permissions"}, "date_modified": "2026-09-24T15:24:17.533329+00:00"}, "schema_version": "1", "submission_centers": [{"uuid": "9626d82e-8110-4213-ac75-0a50adf890ff", "@type": ["SubmissionCenter", "Item"], "status": "open", "@id": "/submission-centers/9626d82e-8110-4213-ac75-0a50adf890ff/", "display_title": "HMS DAC", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "@id": "/static-sections/ce182c42-4f0e-4c1b-b325-213f730a0b42/", "@type": ["StaticSection", "UserContent", "Item"], "uuid": "ce182c42-4f0e-4c1b-b325-213f730a0b42", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}, "display_title": "Illumina (short-read)", "content_as_html": "<div><h2>Overview</h2>\n<p>The paired-end short-read alignment pipeline for Illumina data follows the Genome Analysis Toolkit (GATK) Best Practices. It is designed for per-sample and per-library execution, handling one or multiple sets of paired FASTQ files. The pipeline is optimized for distributed processing, requiring each pair of FASTQ files to correspond to a single sequencing lane.</p>\n<h3>Key Pipeline Steps</h3>\n<ol>\n<li><strong>Alignment with BWA-MEM:</strong> Initial alignment of the raw reads to the reference genome using BWA-MEM.</li>\n<li><strong>Read Groups Assignment:</strong> Assignment of reads to specific groups, providing essential information for subsequent steps.</li>\n<li><strong>Duplicate Reads Marking:</strong> Identification and labeling of duplicate reads, that originated during library preparation or as sequencing artifacts.</li>\n<li><strong>Local Indel Realignment:</strong> Local realignment at identified indel positions.</li>\n<li><strong>Base Quality Score Recalibration (BQSR):</strong> Recalibration of base quality scores, including indels, resulting in the generation of a refined, analysis-ready BAM file.</li>\n</ol>\n<p>A modified alignment step with BWA-MEM is used for processing Hi-C data.</p>\n<h3>Sentieon Software and Distributed Mode</h3>\n<p>To meet the scalability demands, especially with high-coverage whole-genome sequencing data, the pipeline utilizes a more efficient software implementation from <a href=\"https://www.sentieon.com/\" target=\"_blank\" rel=\"noopener noreferrer\">Sentieon</a>.</p>\n<p>Sentieon offers a comprehensive toolkit (DNASeq<sup><sub>1</sub></sup>) that replicates the original BWA and GATK algorithms, while enhancing computational efficiency. The pipeline also leverage Sentieon\u2019s <a href=\"https://support.sentieon.com/appnotes/distributed_mode/\" target=\"_blank\" rel=\"noopener noreferrer\">distributed mode</a> to streamline the processing of large data volumes.</p>\n<p>The distributed version of the pipeline operates on reads divided by sequencing lanes and parallelizing some of the processing across multiple genomic regions, defined as shards. This implementation is a 1 to 1 equivalent to a standard implementation of the pipeline, producing identical results.</p>\n<h3>Pipeline Chart</h3>\n<p><img alt=\"flow_chart\" src=\"/static/img/pipeline-docs/Flow_Chart_Pipeline_Short-Read_Illumina_Paired-End.png\" /></p>\n<p><sub><b>1</b>: <em>Kendig KI, Baheti S, Bockol MA, Drucker TM, Hart SN, Heldenbrand JR, Hernaez M, Hudson ME, Kalmbach MT, Klee EW, Mattson NR, Ross CA, Taschuk M, Wieben ED, Wiepert M, Wildman DE, Mainzer LS.</em> Sentieon DNASeq Variant Calling Workflow Demonstrates Strong Computational Performance and Accuracy. <em>Front Genet. 2019 Aug 20;10:736.</em> d\noi: 10.3389/fgene.2019.00736</sub></p>\n<hr />\n<h2>Alignment with BWA-MEM</h2>\n<p>For the initial alignment, the pipeline uses BWA-MEM in paired-end mode on each set of paired FASTQ files. The reads are then sorted by genomic coordinates, and an integrity check is performed on the resulting BAM file.</p>\n<h3>Aligning and Sorting</h3>\n<h6>Align and sort reads</h6>\n<pre class=\"code-block copy-wrapper\">\nsentieon bwa mem -K 10000000 reference.fasta reads.fastq mates.fastq |\n  samtools sort -o sorted.bam -\n</pre>\n\n<p>Arguments:</p>\n<ul>\n<li><em>-K</em>: chunk size option to have number of threads independent results.</li>\n</ul>\n<h3>Integrity Check</h3>\n<p>To confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.</p>\n<h3>Implementation with Sentieon</h3>\n<p>Sentieon implementation replicates the original <a href=\"https://github.com/lh3/bwa\" target=\"_blank\" rel=\"noopener noreferrer\">BWA</a> code. The pipeline is using Sentieon version 202308.01, corresponding to BWA version 0.7.17.</p>\n<p><em>Note: In the original BWA code there is a programming error affecting alignment score calculations. Specifically, when secondary alignments are present, the score calculation includes unrelated information, potentially influencing the Mapping Quality (MAPQ) of the primary alignment by lowering it. Sentieon introduced a fix to the issue, resulting in slight differences in MAPQ between the two software (approximately 0.008% of reads are affected).</em></p>\n<h3>Source Code</h3>\n<p>All the relevant code can be accessed in the GitHub repository:</p>\n<ul>\n<li><a href=\"https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_bwa-mem_sort.sh\" target=\"_blank\" rel=\"noopener noreferrer\">sentieon_bwa-mem_sort.sh</a> [BWA-MEM]</li>\n</ul>\n<hr />\n<h2>Read Groups</h2>\n<p>As the second step, the pipeline assigns the reads to unique read groups, representing identifiers that group reads together. A read group (<code>@RG</code>) captures relevant information about the sample and the sequencing process and technology, utilized by various downstream bioinformatics tools.</p>\n<p>The relevant fields in defining a read group include:</p>\n<ul>\n<li><strong>ID (Identifier):</strong> A unique identifier for the read group within the BAM file and across multiple BAM files used in the same dataset.</li>\n<li><strong>SM (Sample):</strong> The sample to which the reads belong.</li>\n<li><strong>PL (Platform):</strong> The technology used to sequence the reads. This tag is required if running Base Quality Score Recalibration (BQSR) to determine the correct error model (e.g., ILLUMINA).</li>\n<li><strong>PU (Platform Unit):</strong> A unique identifier for the sequencer unit used for sequencing (i.e., sequencing lane). This tag is required if running BQSR, as it models together all reads belonging to the same platform unit.</li>\n<li><strong>LB (Library):</strong> The library used to sequence the reads. This tag is used in the process of marking or removing duplicate reads to determine groups that may contain duplicates, as duplicate reads need to belong to the same library.</li>\n</ul>\n<h3>Assigning Read Groups</h3>\n<p>To assign read groups, an in-house Python script is used. It can automatically generate read groups based on Illumina read names and handle multiple read groups in the same file (e.g., reads from multiple lanes are merged into a single file).</p>\n<p>The read groups are assigned as follows:</p>\n<ul>\n<li><strong>ID:</strong> <code>&lt;sample name&gt;.&lt;instrument&gt;_&lt;run&gt;_&lt;flow cell&gt;.&lt;lane&gt;</code></li>\n<li><strong>SM:</strong> <code>&lt;sample name&gt;</code></li>\n<li><strong>PL:</strong> <code>&lt;platform&gt;</code></li>\n<li><strong>PU:</strong> <code>&lt;instrument&gt;_&lt;run&gt;_&lt;flow cell&gt;.&lt;lane&gt;</code></li>\n<li><strong>LB:</strong> <code>&lt;sample name&gt;.&lt;library&gt;</code></li>\n</ul>\n<p>E.g., in BAM file:</p>\n<pre class=\"code-block copy-wrapper\">\n@RG ID:SMAHT1.ST-E00127_336_HJ7YHCCXX.8  SM:SMAHT1  PL:ILLUMINA  PU:ST-E00127_336_HJ7YHCCXX.8  LB:SMAHT1.HISEQ-LIB1\n</pre>\n\n<h3>Source Code</h3>\n<p>All the relevant code is accessible in the GitHub repository:</p>\n<ul>\n<li><a href=\"https://github.com/smaht-dac/pipelines-scripts/blob/main/processing_scripts/AddReadGroups.py\" target=\"_blank\" rel=\"noopener noreferrer\">AddReadGroups.py</a> [AddReadGroups]</li>\n</ul>\n<hr />\n<h2>Duplicate Reads</h2>\n<p>In this step, the pipeline marks duplicate reads. Duplicate reads are sequencing artifacts that originate during library preparation and sequencing runs. Duplicate reads are evaluated per-library using the <code>LB</code> tag in the read groups.</p>\n<p>The pipeline does not remove the duplicate reads that are tagged directly in the BAM file.</p>\n<h3>Detecting and Marking Duplicates</h3>\n<h6>Detect duplicate reads</h6>\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -i sorted.bam\n                --algo LocusCollector\n                --fun score_info\n                score.txt\n</pre>\n\n<h6>Mark duplicate reads</h6>\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -i sorted.bam\n                --algo Dedup\n                --optical_dup_pix_dist 2500\n                --score_info score.txt\n                deduped.bam\n</pre>\n\n<p>Arguments:</p>\n<ul>\n<li><em>--optical_dup_pix_dist</em>: maximum offset between two duplicate clusters to consider them optical duplicates. For structured flow cells (NovaSeq, HiSeq 4000, X), the pipeline uses 2500.</li>\n</ul>\n<p><em>Note: The actual implementation of the above command in the pipeline is more complex to support distributed execution, but functionally equivalent.</em></p>\n<h3>Integrity Check</h3>\n<p>To confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.</p>\n<h3>Implementation with Sentieon</h3>\n<p>The pipeline implementation uses Sentieon LocusCollector to calculate duplicate metrics per library and the Dedup algorithm to mark duplicate reads in the BAM file. The pipeline is using Sentieon version 202308.01, corresponding to Picard 2.9.0. Both algorithms combined are equivalent to the MarkDuplicates algorithm in Picard.</p>\n<h6>Detect and mark duplicate reads (Picard equivalent)</h6>\n<pre class=\"code-block copy-wrapper\">\njava -jar picard.jar MarkDuplicates\n      INPUT=sorted.bam\n      OPTICAL_DUPLICATE_PIXEL_DISTANCE=2500\n      OUTPUT=deduped.bam\n</pre>\n\n<h3>Source Code</h3>\n<p>All the relevant code can be accessed in the GitHub repository:</p>\n<ul>\n<li><a href=\"https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_LocusCollector.sh\" target=\"_blank\" rel=\"noopener noreferrer\">sentieon_LocusCollector.sh</a> [LocusCollector]</li>\n<li><a href=\"https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_LocusCollector_apply.sh\" target=\"_blank\" rel=\"noopener noreferrer\">sentieon_LocusCollector_apply.sh</a> [Dedup]</li>\n</ul>\n<hr />\n<h2>Local Realignment</h2>\n<p>In this step, the pipeline perform local realignment at indel positions. Unlike genome aligners that consider each read independently and may favor alignments with mismatches or soft-clips over opening a gap, local realignment takes into account all reads spanning a given position, enabling a high-scoring consensus that supports the presence of an indel event. The two-step realignment process identifies potential regions for alignment improvement and then realigns the reads in these regions using a consensus model that considers all reads in the alignment context together. This allows to correct potential mapping errors made by the aligners, enhancing the consistency of read alignments in regions with indels.</p>\n<h3>Realigning Reads</h3>\n<h6>Realign reads around indels</h6>\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -r reference.fasta\n                -i deduped.bam\n                --algo Realigner\n                -k known_sites_INDEL.vcf\n                realigned.bam\n</pre>\n\n<h3>Integrity Check</h3>\n<p>To confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.</p>\n<h3>Implementation with Sentieon</h3>\n<p>The implementation uses Sentieon Realigner algorithm to identify the potential targets and perform local indel realignment. The pipeline is using Sentieon version 202308.01, corresponding to GATK versions 3.7, 3.8. The algorithm is equivalent to RealignerTargetCreator and IndelRealigner algorithms in GATK.</p>\n<h6>Identify potential targets (GATK equivalent)</h6>\n<pre class=\"code-block copy-wrapper\">\njava -jar GenomeAnalysisTK.jar \n            -T RealignerTargetCreator\n            -R reference.fasta\n            -I deduped.bam \n            -known known_sites_INDEL.vcf\n            -o realignment_targets.list\n</pre>\n\n<h6>Local indel realignment (GATK equivalent)</h6>\n<pre class=\"code-block copy-wrapper\">\njava -jar GenomeAnalysisTK.jar \n            -T IndelRealigner\n            -R reference.fasta\n            -I deduped.bam\n            -targetIntervals realignment_targets.list\n            -known known_sites_INDEL.vcf\n            -o realigned.bam\n</pre>\n\n<h3>Source Code</h3>\n<p>All the relevant code can be accessed in the GitHub repository:</p>\n<ul>\n<li><a href=\"https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_Realigner.sh\" target=\"_blank\" rel=\"noopener noreferrer\">sentieon_Realigner.sh</a> [Realigner]</li>\n</ul>\n<hr />\n<h2>Base Quality Score Recalibration (BQSR)</h2>\n<p>In this final step, the pipeline recalibrates the base quality scores produced by the sequencing machine, correcting systematic (non-random) technical errors leading to over- or under-estimated base quality scores in the data. BQSR modeling is independently performed on groups of reads from different sequencer units (i.e., lanes), identified by the <code>PU</code> tag in the read groups.</p>\n<h3>Base Quality Score Modeling and Recalibration</h3>\n<h6>Model quality scores</h6>\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -r reference.fasta\n                -i realigned.bam\n                --algo QualCal\n                -k known_sites_SNP.vcf\n                -k known_sites_INDEL.vcf\n                recal_data.table\n</pre>\n\n<h6>Apply score recalibration</h6>\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -r reference.fasta\n                -i realigned.bam\n                --read_filter 'QualCalFilter,table=recal_data.table,keep_oq=true'\n                --algo ReadWriter\n                recalibrated.bam\n</pre>\n<p><em>Note: keep_oq=true in the --read_filter argument will preserve the original base quality scores by using the OQ tag in the recalibrated bam file.</em></p>\n<h3>Integrity Check</h3>\n<p>To confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.</p>\n<h3>Implementation with Sentieon</h3>\n<p>The implementation uses Sentieon QualCal algorithm to construct models of covariation based on the input data and a set of known variants. This process produces the recalibration table necessary for BQSR. The recalibration is then applied to the BAM file using the Sentieon ReadWriter command. The pipeline is using Sentieon version 202308.01, corresponding to GATK versions 3.7, 3.8, 4.0, and 4.1. The algorithms are equivalent to BaseRecalibrator and ApplyBQSR algorithms in GATK.</p>\n<h6>Model quality scores (GATK equivalent)</h6>\n<pre class=\"code-block copy-wrapper\">\ngatk BaseRecalibrator -R reference.fasta\n                      -I realigned.bam\n                      --enable-baq\n                      --known-sites known_sites_SNP.vcf\n                      --known-sites known_sites_INDEL.vcf\n                      -O recal_data.table\n</pre>\n\n<h6>Apply score recalibration (GATK equivalent)</h6>\n<pre class=\"code-block copy-wrapper\">\ngatk ApplyBQSR -R reference.fasta\n               -I realigned.bam\n               -bqsr recal_data.table\n               -O recalibrated.bam\n</pre>\n\n<p><em>Note: To reduce run time (at the expense of accuracy), GATK4 disables the base quality score recalibration of indels when using default settings. Consequently, the Sentieon BAM output will contain BI/BD tags from the indel recalibration, which will be missing from the GATK4 BAM output when run with default settings.</em></p>\n<h3>Source Code</h3>\n<p>All the relevant code can be accessed in the GitHub repository:</p>\n<ul>\n<li><a href=\"https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_QualCal.sh\" target=\"_blank\" rel=\"noopener noreferrer\">sentieon_QualCal.sh</a> [QualCal + ReadWriter]</li>\n</ul>\n<hr />\n<h2>Hi-C</h2>\n<p>For the alignment of Hi-C data, the pipeline uses BWA-MEM in paired-end mode on each set of paired FASTQ files. The reads are then sorted by genomic coordinates, and an integrity check is performed on the resulting BAM file. Compared to standard alignment, extra flags are required to support the read pairs generated for the Hi-C data.</p>\n<h3>Aligning and Sorting</h3>\n<h6>Align and sort reads</h6>\n<pre class=\"code-block copy-wrapper\">\nsentieon bwa mem -5SP -K 10000000 reference.fasta reads.fastq mates.fastq |\n  samtools sort -o sorted.bam -\n</pre>\n\n<p>Standard arguments:</p>\n<ul>\n<li><em>-K</em>: chunk size option to have number of threads independent results.</li>\n</ul>\n<p>Hi-C specific arguments:</p>\n<ul>\n<li><em>-S</em>: skip mate rescue.</li>\n<li><em>-P</em>: skip pairing. Mate rescue performed unless -S also in use.</li>\n<li><em>-5</em>: for split alignment, take the alignment with the smallest coordinate as primary.</li>\n</ul>\n<p><em>-SP</em> ensures that the results are equivalent to aligning each mate separately, while maintaining the proper paired-end read formatting. It also avoids forced alignment of a poorly aligned read given an alignment of its mate, based on assumption that the two mates belong to a single genomic segment.</p>\n<p><em>-5</em> ensures that the 5' portion of chimeric alignments is marked as the primary alignment. In Hi-C experiments, the 5' portion of the read is typically the alignment of interest, with the 3' portion representing the same fragment as the mate.</p>\n<h3>Integrity Check</h3>\n<p>To confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.</p>\n<h3>Implementation with Sentieon</h3>\n<p>Sentieon implementation replicates the original <a href=\"https://github.com/lh3/bwa\" target=\"_blank\" rel=\"noopener noreferrer\">BWA</a> code. The pipeline is using Sentieon version 202308.01, corresponding to BWA version 0.7.17.</p>\n<h3>Source Code</h3>\n<p>All the relevant code can be accessed in the GitHub repository:</p>\n<ul>\n<li><a href=\"https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_bwa-mem_sort_Hi-C.sh\" target=\"_blank\" rel=\"noopener noreferrer\">sentieon_bwa-mem_sort_Hi-C.sh</a> [BWA-MEM]</li>\n</ul></div>", "content": "## Overview\n\nThe paired-end short-read alignment pipeline for Illumina data follows the Genome Analysis Toolkit (GATK) Best Practices. It is designed for per-sample and per-library execution, handling one or multiple sets of paired FASTQ files. The pipeline is optimized for distributed processing, requiring each pair of FASTQ files to correspond to a single sequencing lane.\n\n### Key Pipeline Steps\n\n1. **Alignment with BWA-MEM:** Initial alignment of the raw reads to the reference genome using BWA-MEM.\n2. **Read Groups Assignment:** Assignment of reads to specific groups, providing essential information for subsequent steps.\n3. **Duplicate Reads Marking:** Identification and labeling of duplicate reads, that originated during library preparation or as sequencing artifacts.\n4. **Local Indel Realignment:** Local realignment at identified indel positions.\n5. **Base Quality Score Recalibration (BQSR):** Recalibration of base quality scores, including indels, resulting in the generation of a refined, analysis-ready BAM file.\n\nA modified alignment step with BWA-MEM is used for processing Hi-C data.\n\n### Sentieon Software and Distributed Mode\n\nTo meet the scalability demands, especially with high-coverage whole-genome sequencing data, the pipeline utilizes a more efficient software implementation from [Sentieon](https://www.sentieon.com/).\n\nSentieon offers a comprehensive toolkit (DNASeq<sup><sub>1</sub></sup>) that replicates the original BWA and GATK algorithms, while enhancing computational efficiency. The pipeline also leverage Sentieon\u2019s [distributed mode](https://support.sentieon.com/appnotes/distributed_mode/) to streamline the processing of large data volumes.\n\nThe distributed version of the pipeline operates on reads divided by sequencing lanes and parallelizing some of the processing across multiple genomic regions, defined as shards. This implementation is a 1 to 1 equivalent to a standard implementation of the pipeline, producing identical results.\n\n### Pipeline Chart\n\n![flow_chart](/static/img/pipeline-docs/Flow_Chart_Pipeline_Short-Read_Illumina_Paired-End.png)\n\n<sub><b>1</b>: *Kendig KI, Baheti S, Bockol MA, Drucker TM, Hart SN, Heldenbrand JR, Hernaez M, Hudson ME, Kalmbach MT, Klee EW, Mattson NR, Ross CA, Taschuk M, Wieben ED, Wiepert M, Wildman DE, Mainzer LS.* Sentieon DNASeq Variant Calling Workflow Demonstrates Strong Computational Performance and Accuracy. *Front Genet. 2019 Aug 20;10:736.* d\noi: 10.3389/fgene.2019.00736</sub>\n\n---\n\n## Alignment with BWA-MEM\n\nFor the initial alignment, the pipeline uses BWA-MEM in paired-end mode on each set of paired FASTQ files. The reads are then sorted by genomic coordinates, and an integrity check is performed on the resulting BAM file.\n\n### Aligning and Sorting\n\n###### Align and sort reads\n\n<pre class=\"code-block copy-wrapper\">\nsentieon bwa mem -K 10000000 reference.fasta reads.fastq mates.fastq |\n  samtools sort -o sorted.bam -\n</pre>\n\nArguments:\n\n- *-K*: chunk size option to have number of threads independent results.\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nSentieon implementation replicates the original [BWA](https://github.com/lh3/bwa) code. The pipeline is using Sentieon version 202308.01, corresponding to BWA version 0.7.17.\n\n*Note: In the original BWA code there is a programming error affecting alignment score calculations. Specifically, when secondary alignments are present, the score calculation includes unrelated information, potentially influencing the Mapping Quality (MAPQ) of the primary alignment by lowering it. Sentieon introduced a fix to the issue, resulting in slight differences in MAPQ between the two software (approximately 0.008% of reads are affected).*\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_bwa-mem_sort.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_bwa-mem_sort.sh) [BWA-MEM]\n\n---\n\n## Read Groups\n\nAs the second step, the pipeline assigns the reads to unique read groups, representing identifiers that group reads together. A read group (`@RG`) captures relevant information about the sample and the sequencing process and technology, utilized by various downstream bioinformatics tools.\n\nThe relevant fields in defining a read group include:\n\n- **ID (Identifier):** A unique identifier for the read group within the BAM file and across multiple BAM files used in the same dataset.\n- **SM (Sample):** The sample to which the reads belong.\n- **PL (Platform):** The technology used to sequence the reads. This tag is required if running Base Quality Score Recalibration (BQSR) to determine the correct error model (e.g., ILLUMINA).\n- **PU (Platform Unit):** A unique identifier for the sequencer unit used for sequencing (i.e., sequencing lane). This tag is required if running BQSR, as it models together all reads belonging to the same platform unit.\n- **LB (Library):** The library used to sequence the reads. This tag is used in the process of marking or removing duplicate reads to determine groups that may contain duplicates, as duplicate reads need to belong to the same library.\n\n### Assigning Read Groups\n\nTo assign read groups, an in-house Python script is used. It can automatically generate read groups based on Illumina read names and handle multiple read groups in the same file (e.g., reads from multiple lanes are merged into a single file).\n\nThe read groups are assigned as follows:\n\n- **ID:** `<sample name>.<instrument>_<run>_<flow cell>.<lane>`\n- **SM:** `<sample name>`\n- **PL:** `<platform>`\n- **PU:** `<instrument>_<run>_<flow cell>.<lane>`\n- **LB:** `<sample name>.<library>`\n\nE.g., in BAM file:\n\n<pre class=\"code-block copy-wrapper\">\n@RG ID:SMAHT1.ST-E00127_336_HJ7YHCCXX.8  SM:SMAHT1  PL:ILLUMINA  PU:ST-E00127_336_HJ7YHCCXX.8  LB:SMAHT1.HISEQ-LIB1\n</pre>\n\n### Source Code\n\nAll the relevant code is accessible in the GitHub repository:\n\n  - [AddReadGroups.py](https://github.com/smaht-dac/pipelines-scripts/blob/main/processing_scripts/AddReadGroups.py) [AddReadGroups]\n\n---\n\n## Duplicate Reads\n\nIn this step, the pipeline marks duplicate reads. Duplicate reads are sequencing artifacts that originate during library preparation and sequencing runs. Duplicate reads are evaluated per-library using the `LB` tag in the read groups.\n\nThe pipeline does not remove the duplicate reads that are tagged directly in the BAM file.\n\n### Detecting and Marking Duplicates\n\n###### Detect duplicate reads\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -i sorted.bam\n                --algo LocusCollector\n                --fun score_info\n                score.txt\n</pre>\n\n###### Mark duplicate reads\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -i sorted.bam\n                --algo Dedup\n                --optical_dup_pix_dist 2500\n                --score_info score.txt\n                deduped.bam\n</pre>\n\nArguments:\n\n- *-\\-optical_dup_pix_dist*: maximum offset between two duplicate clusters to consider them optical duplicates. For structured flow cells (NovaSeq, HiSeq 4000, X), the pipeline uses 2500.\n\n*Note: The actual implementation of the above command in the pipeline is more complex to support distributed execution, but functionally equivalent.*\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nThe pipeline implementation uses Sentieon LocusCollector to calculate duplicate metrics per library and the Dedup algorithm to mark duplicate reads in the BAM file. The pipeline is using Sentieon version 202308.01, corresponding to Picard 2.9.0. Both algorithms combined are equivalent to the MarkDuplicates algorithm in Picard.\n\n###### Detect and mark duplicate reads (Picard equivalent)\n\n<pre class=\"code-block copy-wrapper\">\njava -jar picard.jar MarkDuplicates\n      INPUT=sorted.bam\n      OPTICAL_DUPLICATE_PIXEL_DISTANCE=2500\n      OUTPUT=deduped.bam\n</pre>\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_LocusCollector.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_LocusCollector.sh) [LocusCollector]\n  - [sentieon_LocusCollector_apply.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_LocusCollector_apply.sh) [Dedup]\n\n---\n\n## Local Realignment\n\nIn this step, the pipeline perform local realignment at indel positions. Unlike genome aligners that consider each read independently and may favor alignments with mismatches or soft-clips over opening a gap, local realignment takes into account all reads spanning a given position, enabling a high-scoring consensus that supports the presence of an indel event. The two-step realignment process identifies potential regions for alignment improvement and then realigns the reads in these regions using a consensus model that considers all reads in the alignment context together. This allows to correct potential mapping errors made by the aligners, enhancing the consistency of read alignments in regions with indels.\n\n### Realigning Reads\n\n###### Realign reads around indels\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -r reference.fasta\n                -i deduped.bam\n                --algo Realigner\n                -k known_sites_INDEL.vcf\n                realigned.bam\n</pre>\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nThe implementation uses Sentieon Realigner algorithm to identify the potential targets and perform local indel realignment. The pipeline is using Sentieon version 202308.01, corresponding to GATK versions 3.7, 3.8. The algorithm is equivalent to RealignerTargetCreator and IndelRealigner algorithms in GATK.\n\n###### Identify potential targets (GATK equivalent)\n\n<pre class=\"code-block copy-wrapper\">\njava -jar GenomeAnalysisTK.jar \n            -T RealignerTargetCreator\n            -R reference.fasta\n            -I deduped.bam \n            -known known_sites_INDEL.vcf\n            -o realignment_targets.list\n</pre>\n\n###### Local indel realignment (GATK equivalent)\n\n<pre class=\"code-block copy-wrapper\">\njava -jar GenomeAnalysisTK.jar \n            -T IndelRealigner\n            -R reference.fasta\n            -I deduped.bam\n            -targetIntervals realignment_targets.list\n            -known known_sites_INDEL.vcf\n            -o realigned.bam\n</pre>\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_Realigner.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_Realigner.sh) [Realigner]\n\n---\n\n## Base Quality Score Recalibration (BQSR)\n\nIn this final step, the pipeline recalibrates the base quality scores produced by the sequencing machine, correcting systematic (non-random) technical errors leading to over- or under-estimated base quality scores in the data. BQSR modeling is independently performed on groups of reads from different sequencer units (i.e., lanes), identified by the `PU` tag in the read groups.\n\n### Base Quality Score Modeling and Recalibration\n\n###### Model quality scores\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -r reference.fasta\n                -i realigned.bam\n                --algo QualCal\n                -k known_sites_SNP.vcf\n                -k known_sites_INDEL.vcf\n                recal_data.table\n</pre>\n\n###### Apply score recalibration\n\n<pre class=\"code-block copy-wrapper\">\nsentieon driver -r reference.fasta\n                -i realigned.bam\n                --read_filter 'QualCalFilter,table=recal_data.table,keep_oq=true'\n                --algo ReadWriter\n                recalibrated.bam\n</pre>\n*Note: keep_oq=true in the -\\-read_filter argument will preserve the original base quality scores by using the OQ tag in the recalibrated bam file.*\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nThe implementation uses Sentieon QualCal algorithm to construct models of covariation based on the input data and a set of known variants. This process produces the recalibration table necessary for BQSR. The recalibration is then applied to the BAM file using the Sentieon ReadWriter command. The pipeline is using Sentieon version 202308.01, corresponding to GATK versions 3.7, 3.8, 4.0, and 4.1. The algorithms are equivalent to BaseRecalibrator and ApplyBQSR algorithms in GATK.\n\n###### Model quality scores (GATK equivalent)\n\n<pre class=\"code-block copy-wrapper\">\ngatk BaseRecalibrator -R reference.fasta\n                      -I realigned.bam\n                      --enable-baq\n                      --known-sites known_sites_SNP.vcf\n                      --known-sites known_sites_INDEL.vcf\n                      -O recal_data.table\n</pre>\n\n###### Apply score recalibration (GATK equivalent)\n\n<pre class=\"code-block copy-wrapper\">\ngatk ApplyBQSR -R reference.fasta\n               -I realigned.bam\n               -bqsr recal_data.table\n               -O recalibrated.bam\n</pre>\n\n*Note: To reduce run time (at the expense of accuracy), GATK4 disables the base quality score recalibration of indels when using default settings. Consequently, the Sentieon BAM output will contain BI/BD tags from the indel recalibration, which will be missing from the GATK4 BAM output when run with default settings.*\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_QualCal.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_QualCal.sh) [QualCal + ReadWriter]\n\n---\n\n## Hi-C\n\nFor the alignment of Hi-C data, the pipeline uses BWA-MEM in paired-end mode on each set of paired FASTQ files. The reads are then sorted by genomic coordinates, and an integrity check is performed on the resulting BAM file. Compared to standard alignment, extra flags are required to support the read pairs generated for the Hi-C data.\n\n### Aligning and Sorting\n\n###### Align and sort reads\n\n<pre class=\"code-block copy-wrapper\">\nsentieon bwa mem -5SP -K 10000000 reference.fasta reads.fastq mates.fastq |\n  samtools sort -o sorted.bam -\n</pre>\n\nStandard arguments:\n\n- *-K*: chunk size option to have number of threads independent results.\n\nHi-C specific arguments:\n\n- *-S*: skip mate rescue.\n- *-P*: skip pairing. Mate rescue performed unless -S also in use.\n- *-5*: for split alignment, take the alignment with the smallest coordinate as primary.\n\n*-SP* ensures that the results are equivalent to aligning each mate separately, while maintaining the proper paired-end read formatting. It also avoids forced alignment of a poorly aligned read given an alignment of its mate, based on assumption that the two mates belong to a single genomic segment.\n\n*-5* ensures that the 5' portion of chimeric alignments is marked as the primary alignment. In Hi-C experiments, the 5' portion of the read is typically the alignment of interest, with the 3' portion representing the same fragment as the mate.\n\n### Integrity Check\n\nTo confirm the integrity of the alignment BAM file, in-house Python code checks for the presence of the 28-byte empty block representing the EOF marker in BAM format.\n\n### Implementation with Sentieon\n\nSentieon implementation replicates the original [BWA](https://github.com/lh3/bwa) code. The pipeline is using Sentieon version 202308.01, corresponding to BWA version 0.7.17.\n\n### Source Code\n\nAll the relevant code can be accessed in the GitHub repository:\n\n  - [sentieon_bwa-mem_sort_Hi-C.sh](https://github.com/smaht-dac/sentieon-pipelines/blob/main/dockerfiles/sentieon/sentieon_bwa-mem_sort_Hi-C.sh) [BWA-MEM]", "filetype": "md"}], "consortia": [{"uuid": "358aed10-9b9d-4e26-ab84-4bd162da182b", "display_title": "SMaHT", "@id": "/consortia/358aed10-9b9d-4e26-ab84-4bd162da182b/", "@type": ["Consortium", "Item"], "status": "open", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "identifier": "docs/additional-resources/pipeline-docs/short-read_illumina_paired-end", "date_created": "2026-01-09T20:07:38.467642+00:00", "submitted_by": {"error": "no view permissions"}, "last_modified": {"modified_by": {"error": "no view permissions"}, "date_modified": "2026-09-24T15:25:02.335218+00:00"}, "schema_version": "1", "table-of-contents": {"enabled": true, "skip-depth": 1, "header-depth": 2, "include-top-link": false}, "submission_centers": [{"@type": ["SubmissionCenter", "Item"], "display_title": "HMS DAC", "status": "open", "@id": "/submission-centers/9626d82e-8110-4213-ac75-0a50adf890ff/", "uuid": "9626d82e-8110-4213-ac75-0a50adf890ff", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "@id": "/docs/additional-resources/pipeline-docs/short-read_illumina_paired-end", "@type": ["DocsAdditional-resourcesPipeline-docsShort-read_illumina_paired-endPage", "DocsAdditional-resourcesPipeline-docsPage", "DocsAdditional-resourcesPage", "DocsPage", "StaticPage", "Portal"], "uuid": "1756229d-7498-444f-b574-983d3020f948", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}, "display_title": "Illumina (short-read)", "@context": "/docs/additional-resources/pipeline-docs/short-read_illumina_paired-end", "is_leaf": true, "toc": {"enabled": true, "skip-depth": 1, "header-depth": 2, "include-top-link": false}, "next": {"identifier": "docs/additional-resources/pipeline-docs/long-read_pacbio_hifi", "title": "PacBio HiFi (long-read)", "status": "open", "content": [{"status": "open", "uuid": "23bd319b-19eb-4408-869b-eda806da56a5", "@type": ["StaticSection", "UserContent", "Item"], "display_title": "PacBio HiFi (long-read)", "@id": "/static-sections/23bd319b-19eb-4408-869b-eda806da56a5/", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "consortia": [{"@id": "/consortia/358aed10-9b9d-4e26-ab84-4bd162da182b/", "status": "open", "uuid": "358aed10-9b9d-4e26-ab84-4bd162da182b", "@type": ["Consortium", "Item"], "display_title": "SMaHT", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "date_created": "2026-01-09T20:07:38.631372+00:00", "submitted_by": {"error": "no view permissions"}, "last_modified": {"modified_by": {"error": "no view permissions"}, "date_modified": "2026-09-24T15:25:02.567073+00:00"}, "schema_version": "1", "table-of-contents": {"enabled": true, "skip-depth": 1, "header-depth": 2, "include-top-link": false}, "submission_centers": [{"uuid": "9626d82e-8110-4213-ac75-0a50adf890ff", "display_title": "HMS DAC", "@type": ["SubmissionCenter", "Item"], "@id": "/submission-centers/9626d82e-8110-4213-ac75-0a50adf890ff/", "status": "open", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "@id": "/docs/additional-resources/pipeline-docs/long-read_pacbio_hifi", "@type": ["DocsAdditional-resourcesPipeline-docsLong-read_pacbio_hifiPage", "DocsAdditional-resourcesPipeline-docsPage", "DocsAdditional-resourcesPage", "DocsPage", "StaticPage", "Portal"], "uuid": "ab58ce07-b8f9-4807-aa8a-d651b6cd733d", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}, "display_title": "PacBio HiFi (long-read)", "is_leaf": true, "sibling_length": 13, "sibling_position": 2}, "previous": {"identifier": "docs/additional-resources/pipeline-docs/fastq_files", "title": "FASTQ", "status": "open", "content": [{"status": "open", "uuid": "439034da-bad5-42c5-8124-3be68276d6cc", "@type": ["StaticSection", "UserContent", "Item"], "display_title": "FASTQ", "@id": "/static-sections/439034da-bad5-42c5-8124-3be68276d6cc/", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "consortia": [{"@id": "/consortia/358aed10-9b9d-4e26-ab84-4bd162da182b/", "status": "open", "uuid": "358aed10-9b9d-4e26-ab84-4bd162da182b", "@type": ["Consortium", "Item"], "display_title": "SMaHT", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "date_created": "2026-01-09T20:07:38.300564+00:00", "submitted_by": {"error": "no view permissions"}, "last_modified": {"modified_by": {"error": "no view permissions"}, "date_modified": "2026-09-24T15:25:02.142170+00:00"}, "schema_version": "1", "table-of-contents": {"enabled": true, "skip-depth": 1, "header-depth": 2, "include-top-link": false}, "submission_centers": [{"uuid": "9626d82e-8110-4213-ac75-0a50adf890ff", "display_title": "HMS DAC", "@type": ["SubmissionCenter", "Item"], "@id": "/submission-centers/9626d82e-8110-4213-ac75-0a50adf890ff/", "status": "open", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "@id": "/docs/additional-resources/pipeline-docs/fastq_files", "@type": ["DocsAdditional-resourcesPipeline-docsFastq_filesPage", "DocsAdditional-resourcesPipeline-docsPage", "DocsAdditional-resourcesPage", "DocsPage", "StaticPage", "Portal"], "uuid": "379e3032-a0fe-45d6-8277-b373b52d3a79", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}, "display_title": "FASTQ", "is_leaf": true, "sibling_length": 13, "sibling_position": 0}, "parent": {"identifier": "docs/additional-resources/pipeline-docs", "parent": {"identifier": "docs/additional-resources", "parent": {"identifier": "docs", "parent": {"identifier": "", "@id": "/", "display_title": "Home", "@type": ["DirectoryPage", "StaticPage", "Portal"]}, "@id": "/docs", "uuid": "089319c4-3ce9-4ec1-bd0b-5451a48bd99e", "display_title": "Documentation", "@type": ["DocsPage", "DirectoryPage", "StaticPage", "Portal"], "sibling_length": 6, "sibling_position": 3}, "title": "Analysis & Additional Resources", "status": "open", "consortia": [{"@type": ["Consortium", "Item"], "@id": "/consortia/358aed10-9b9d-4e26-ab84-4bd162da182b/", "display_title": "SMaHT", "uuid": "358aed10-9b9d-4e26-ab84-4bd162da182b", "status": "open", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "date_created": "2024-03-01T19:21:24.278212+00:00", "submitted_by": {"error": "no view permissions"}, "last_modified": {"modified_by": {"error": "no view permissions"}, "date_modified": "2026-09-24T15:25:07.179445+00:00"}, "schema_version": "1", "table-of-contents": {"enabled": true, "skip-depth": 1, "header-depth": 4, "include-top-link": false}, "submission_centers": [{"status": "open", "display_title": "HMS DAC", "uuid": "9626d82e-8110-4213-ac75-0a50adf890ff", "@id": "/submission-centers/9626d82e-8110-4213-ac75-0a50adf890ff/", "@type": ["SubmissionCenter", "Item"], "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "@id": "/docs/additional-resources", "@type": ["DocsAdditional-resourcesPage", "DocsPage", "DirectoryPage", "StaticPage", "Portal"], "uuid": "1ada4fca-af4b-4304-947d-59e2918ab728", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}, "display_title": "Analysis & Additional Resources", "sibling_length": 3, "sibling_position": 2}, "title": "Analysis Pipelines", "status": "open", "content": [{"@type": ["StaticSection", "UserContent", "Item"], "uuid": "b78b2ebb-d01c-4635-8a67-76a98ab81772", "status": "open", "@id": "/static-sections/b78b2ebb-d01c-4635-8a67-76a98ab81772/", "display_title": "Analysis Pipelines", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "redirect": {"code": 307, "enabled": false}, "consortia": [{"status": "open", "display_title": "SMaHT", "@type": ["Consortium", "Item"], "@id": "/consortia/358aed10-9b9d-4e26-ab84-4bd162da182b/", "uuid": "358aed10-9b9d-4e26-ab84-4bd162da182b", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "date_created": "2026-01-09T20:07:37.883644+00:00", "submitted_by": {"error": "no view permissions"}, "last_modified": {"modified_by": {"error": "no view permissions"}, "date_modified": "2026-09-24T15:25:01.829129+00:00"}, "schema_version": "1", "submission_centers": [{"uuid": "9626d82e-8110-4213-ac75-0a50adf890ff", "@id": "/submission-centers/9626d82e-8110-4213-ac75-0a50adf890ff/", "@type": ["SubmissionCenter", "Item"], "status": "open", "display_title": "HMS DAC", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}}], "@id": "/docs/additional-resources/pipeline-docs", "@type": ["DocsAdditional-resourcesPipeline-docsPage", "DocsAdditional-resourcesPage", "DocsPage", "DirectoryPage", "StaticPage", "Portal"], "uuid": "6e144832-6abc-47e2-bea5-f720598cf61a", "principals_allowed": {"view": ["system.Everyone"], "edit": ["group.admin"]}, "display_title": "Analysis Pipelines", "sibling_length": 5, "sibling_position": 0}, "sibling_length": 13, "sibling_position": 1}