<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://en.formulasearchengine.com/w/index.php?action=history&amp;feed=atom&amp;title=Glossary_of_stack_theory</id>
	<title>Glossary of stack theory - Revision history</title>
	<link rel="self" type="application/atom+xml" href="https://en.formulasearchengine.com/w/index.php?action=history&amp;feed=atom&amp;title=Glossary_of_stack_theory"/>
	<link rel="alternate" type="text/html" href="https://en.formulasearchengine.com/w/index.php?title=Glossary_of_stack_theory&amp;action=history"/>
	<updated>2026-07-31T03:50:31Z</updated>
	<subtitle>Revision history for this page on the wiki</subtitle>
	<generator>MediaWiki 1.47.0-wmf.7</generator>
	<entry>
		<id>https://en.formulasearchengine.com/w/index.php?title=Glossary_of_stack_theory&amp;diff=30106&amp;oldid=prev</id>
		<title>en&gt;Malcolma: added Category:Category theory; removed {{uncategorized}} using HotCat</title>
		<link rel="alternate" type="text/html" href="https://en.formulasearchengine.com/w/index.php?title=Glossary_of_stack_theory&amp;diff=30106&amp;oldid=prev"/>
		<updated>2013-12-08T11:08:54Z</updated>

		<summary type="html">&lt;p&gt;added &lt;a href=&quot;/w/index.php?title=Category:Category_theory&amp;amp;action=edit&amp;amp;redlink=1&quot; class=&quot;new&quot; title=&quot;Category:Category theory (page does not exist)&quot;&gt;Category:Category theory&lt;/a&gt;; removed {{uncategorized}} using &lt;a href=&quot;/w/index.php?title=WP:HC&amp;amp;action=edit&amp;amp;redlink=1&quot; class=&quot;new&quot; title=&quot;WP:HC (page does not exist)&quot;&gt;HotCat&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;b&gt;New page&lt;/b&gt;&lt;/p&gt;&lt;div&gt;{{Orphan|date=October 2013}}&lt;br /&gt;
&lt;br /&gt;
&amp;#039;&amp;#039;&amp;#039;SNV calling from NGS data&amp;#039;&amp;#039;&amp;#039; refers to a range of methods for identifying the existence of [[single nucleotide polymorphism|single nucleotide variants]] (SNVs) from the results of [[Next generation sequencing|NGS]] experiments. These are computational techniques, and are in contrast to specialist experimental methods based on known population-wide single nucleotide polymorphisms (see [[SNP genotyping]]). Due to the increasing abundance of NGS data, these techniques are becoming increasingly popular for performing SNP genotyping, with a wide variety of algorithms designed for specific experimental designs and applications.&amp;lt;ref name = &amp;quot;Nielsen 2011&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=Genotype and SNP calling from next-generation sequencing data&lt;br /&gt;
  |author=Nielsen, Rasmus and Paul, Joshua S and Albrechtsen, Anders and Song, Yun S&lt;br /&gt;
  |journal=Nature Reviews Genetics&lt;br /&gt;
  |volume=12&lt;br /&gt;
  |number=6&lt;br /&gt;
  |pages=443–451&lt;br /&gt;
  |year=2011&lt;br /&gt;
  |publisher=Nature Publishing Group&lt;br /&gt;
}}&amp;lt;/ref&amp;gt; In addition to the usual application domain of SNP genotyping, these techniques have been successfully adapted to identify rare SNPs within a population,&amp;lt;ref name = &amp;quot;Bansal 2010&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=A statistical method for the detection of variants from next-generation resequencing of DNA pools&lt;br /&gt;
  |author=Bansal, Vikas&lt;br /&gt;
  |journal=Bioinformatics&lt;br /&gt;
  |volume=26&lt;br /&gt;
  |number=12&lt;br /&gt;
  |pages=i318-i324&lt;br /&gt;
  |year=2010&lt;br /&gt;
  |publisher=Oxford Univ Press&lt;br /&gt;
}}&amp;lt;/ref&amp;gt; as well as detecting [[somatic]] SNVs within an individual using multiple tissue samples.&amp;lt;ref name = &amp;quot;Roth 2012&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=JointSNVMix: a probabilistic model for accurate detection of[somatic mutations in normal/tumour paired next-generation sequencing data&lt;br /&gt;
  |author=Roth, Andrew and Ding, Jiarui and Morin, Ryan and Crisan, Anamaria and Ha, Gavin and Giuliany, Ryan and Bashashati, Ali and Hirst, Martin and Turashvili, Gulisa and Oloumi, Arusha and others&lt;br /&gt;
  |journal=Bioinformatics&lt;br /&gt;
  |volume=28&lt;br /&gt;
  |number=7&lt;br /&gt;
  |pages=907–913&lt;br /&gt;
  |year=2012&lt;br /&gt;
  |publisher=Oxford Univ Press&lt;br /&gt;
}}&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Methods for detecting germline variants ==&lt;br /&gt;
&lt;br /&gt;
Most NGS based methods for SNV detection are designed to detect [[germline]] variations in the individual&amp;#039;s genome. These are the mutations that an individual biologically inherits from their parents, and are the usual type of variants searched for when performing such analysis (except for certain specific applications where [[#Methods for detecting somatic variants|somatic mutations]] are sought). Very often, the searched for variants occur with some (possibly rare) frequency, throughout the population, in which case they may be referred to as [[single nucleotide polymorphisms]] (SNPs). Technically the term SNP only refers to these kinds of variations, however in practice they are often used synonymously with SNV in the literature on variant calling. In addition, since the detection of germline SNVs requires determining the individual&amp;#039;s genotype at each locus, the phrase &amp;quot;SNP genotyping&amp;quot; may also be used to refer to this process. However this phrase may also refer to wet-lab experimental procedures for classifying genotypes at a set of known SNP locations.&lt;br /&gt;
&lt;br /&gt;
The usual process of such techniques are based around:&amp;lt;ref name=&amp;quot;Nielsen 2011&amp;quot;/&amp;gt;&lt;br /&gt;
# Filtering the set of NGS reads to remove sources of error/bias&lt;br /&gt;
# Aligning the reads to a reference genome&lt;br /&gt;
# Using an algorithm, either based on a statistical model or some heuristics, to predict the likelihood of variation at each locus, based on the quality scores and allele counts of the aligned reads at that locus&lt;br /&gt;
# Filtering the predicted results, often based on metrics relevant to the application&lt;br /&gt;
The usual output of these procedures is a [[Variant Call Format|VCF]] file.&lt;br /&gt;
&lt;br /&gt;
=== Probabilistic methods ===&lt;br /&gt;
&lt;br /&gt;
[[File:Heterozygous SNV call, from aligned NGS reads.png|thumb|right|A set of hypothetical NGS reads are shown, aligned against a reference sequence. At the annotated locus, the reads contain a mixture of A/G nucleotides, against the A reference allele. Depending on the prior genotype probabilities, and the chosen error model, this may be called as a heterozygous SNV (genotype AG predicted), the G nucleotides may be classified as errors and no variant called (genotype AA predicted), or alternatively the A nucleotides may be classified as errors and a homozygous SNV valled (genotype GG predicted).]]&lt;br /&gt;
&lt;br /&gt;
In an ideal error free world with high read coverage, the task of variant calling from the results of a NGS data alignment would be simple; at each [[Locus (genetics)|locus]] (position on the genome) the number of occurrences of each distinct nucleotide among the reads alinged at that position can be counted, and the true genotype would be obvious; either &amp;#039;&amp;#039;&amp;#039;AA&amp;#039;&amp;#039;&amp;#039; if all nucleotides match allele &amp;#039;&amp;#039;&amp;#039;A&amp;#039;&amp;#039;&amp;#039;, &amp;#039;&amp;#039;&amp;#039;BB&amp;#039;&amp;#039;&amp;#039; if they match allele &amp;#039;&amp;#039;&amp;#039;B&amp;#039;&amp;#039;&amp;#039;, or &amp;#039;&amp;#039;&amp;#039;AB&amp;#039;&amp;#039;&amp;#039; if there is a mixture. However when working with real NGS data this sort of naive approach is not used, as it cannot account for the noise in the input data.&amp;lt;ref name=&amp;quot;Martin 2010&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=SeqEM: an adaptive genotype-calling approach for next-generation sequencing studies&lt;br /&gt;
  |author=Martin, Eden R and Kinnamon, DD and Schmidt, Michael A and Powell, EH and Zuchner, S and Morris, RW&lt;br /&gt;
  |journal=Bioinformatics&lt;br /&gt;
  |volume=26&lt;br /&gt;
  |number=22&lt;br /&gt;
  |pages=2803–2810&lt;br /&gt;
  |year=2010&lt;br /&gt;
  |publisher=Oxford Univ Press&lt;br /&gt;
}}&amp;lt;/ref&amp;gt; The nucleotide counts used for base calling contain errors and bias, both due do the sequenced reads themselves, and the alignment process. This issue can be mitigated to some extent by sequencing  to a greater depth of read coverage, however this is often expensive, and many practical studies require making inferences on low coverage data.&amp;lt;ref name=&amp;quot;Nielsen 2011&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Probabilistic methods aim to overcome the above issue, by producing robust estimates of the probabilities of each of the possible genotypes, taking into account noise, as well as other available prior information that can be used to improve estimates. A genotype can then be predicted based on these probabilities, often according to the [[Maximum a posteriori estimation|MAP]] estimate.&lt;br /&gt;
&lt;br /&gt;
Probabilistic methods for variant calling are based on [[Bayes Theorem|Bayes&amp;#039; Theorem]]. In the context of variant calling, Bayes&amp;#039; Theorem defines the probability of each genotype being the true genotype given the observed data, in terms of the prior probabilities of each possible genotype, and the probability distribution of the data given each possible genotype. The formula is:&lt;br /&gt;
:&amp;lt;math&amp;gt;&lt;br /&gt;
\begin{align}&lt;br /&gt;
P(G|D) &amp;amp;=  \frac{P(D|G) P(G)}{P(D)}\\[8pt]&lt;br /&gt;
&amp;amp;= \frac{P(D|G)\,P(G)}{\sum\limits_{i=1}^{n} P(D|G_i)\,P(G_i)}\\[8pt]&lt;br /&gt;
\end{align}&amp;lt;/math&amp;gt;&lt;br /&gt;
In the above equation:&lt;br /&gt;
*&amp;lt;math&amp;gt;D&amp;lt;/math&amp;gt; refers to the observed data; that is, the aligned reads&lt;br /&gt;
*&amp;lt;math&amp;gt;G&amp;lt;/math&amp;gt; is the genotype whose probability is being calculated&lt;br /&gt;
*&amp;lt;math&amp;gt;G_i&amp;lt;/math&amp;gt; refers to the &amp;#039;&amp;#039;i&amp;#039;&amp;#039;th possible genotype, out of &amp;#039;&amp;#039;n&amp;#039;&amp;#039; possibilities&lt;br /&gt;
&lt;br /&gt;
Given the above framework, different software solutions for detecting SNVs vary based on how they calculate the prior probabilities &amp;lt;math&amp;gt;P(G)&amp;lt;/math&amp;gt;, the error model used to model the probabilities &amp;lt;math&amp;gt;P(D|G)&amp;lt;/math&amp;gt;, and the partitioning of the overall genotypes into separate sub-genotypes, whose probabilities can be individually estimated in this framework.&amp;lt;ref name=&amp;quot;You 2012&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=SNP calling using genotype model selection on high-throughput sequencing data&lt;br /&gt;
  |author=You, Na and Murillo, Gabriel and Su, Xiaoquan and Zeng, Xiaowei and Xu, Jian and Ning, Kang and Zhang, Shoudong and Zhu, Jiankang and Cui, Xinping&lt;br /&gt;
  |journal=Bioinformatics&lt;br /&gt;
  |volume=28&lt;br /&gt;
  |number=5&lt;br /&gt;
  |pages=643–650&lt;br /&gt;
  |year=2012&lt;br /&gt;
  |publisher=Oxford Univ Press&lt;br /&gt;
}}&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Prior genotype probability estimation ====&lt;br /&gt;
&lt;br /&gt;
The calculation of prior probabilities depends on available data from the genome being studied, and the type of analysis being performed. For studies where good reference data containing frequencies of known mutations is available (for example, in studying human genome data), these known frequencies of genotypes in the population can be used to estimate priors. Given population wide allele frequencies, prior genotype probabilities can be calculated at each locus according to the [[Hardy Weinberg Equilibrium]].&amp;lt;ref name =&amp;quot;Li 2009&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=SNP detection for massively parallel whole-genome resequencing&lt;br /&gt;
  |author=Li, Ruiqiang and Li, Yingrui and Fang, Xiaodong and Yang, Huanming and Wang, Jian and Kristiansen, Karsten and Wang, Jun&lt;br /&gt;
  |journal=Genome research&lt;br /&gt;
  |volume=19&lt;br /&gt;
  |number=6&lt;br /&gt;
  |pages=1124–1132&lt;br /&gt;
  |year=2009&lt;br /&gt;
  |publisher=Cold Spring Harbor Lab&lt;br /&gt;
}}&amp;lt;/ref&amp;gt; In the absence of such data, constant priors can be used, independent of the locus. These can be set using heuristically chosen values, possibly informed by the kind of variations being sought by the study. Alternatively, supervised machine-learning procedures have been investigated that seek to learn optimal prior values for individuals in a sample, using supplied NGS data from these individuals.&amp;lt;ref name=&amp;quot;Martin 2010&amp;quot;&amp;gt;{{cite jounral&lt;br /&gt;
  |title=SeqEM: an adaptive genotype-calling approach for next-generation sequencing studies&lt;br /&gt;
  |author=Martin, Eden R and Kinnamon, DD and Schmidt, Michael A and Powell, EH and Zuchner, S and Morris, RW&lt;br /&gt;
  |journal=Bioinformatics&lt;br /&gt;
  |volume=26&lt;br /&gt;
  |number=22&lt;br /&gt;
  |pages=2803-2810&lt;br /&gt;
  |year=2010&lt;br /&gt;
  |publisher=Oxford Univ Press&lt;br /&gt;
}}&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Error models for data observations ====&lt;br /&gt;
&lt;br /&gt;
The error model used in creating a probabilistic method for variant calling is the basis for calculating the &amp;lt;math&amp;gt;P(D|G)&amp;lt;/math&amp;gt; term used in Bayes&amp;#039; theorem. If the data was assumed to be error free, then the distribution of observed nucleotide counts at each locus would follow a [[Binomial Distribution]], with 100% of nucleotides matching the A or B allele respectively in the &amp;#039;&amp;#039;&amp;#039;AA&amp;#039;&amp;#039;&amp;#039; and &amp;#039;&amp;#039;&amp;#039;BB&amp;#039;&amp;#039;&amp;#039; cases, and a 50% chance of each nucleotide matching either &amp;#039;&amp;#039;&amp;#039;A&amp;#039;&amp;#039;&amp;#039; or &amp;#039;&amp;#039;&amp;#039;B&amp;#039;&amp;#039;&amp;#039; in the &amp;#039;&amp;#039;&amp;#039;AB&amp;#039;&amp;#039;&amp;#039; case. However in presence of noise in the read data this assumption is violated, and the &amp;lt;math&amp;gt;P(D|G)&amp;lt;/math&amp;gt; values need to account for the possibility that erroneous nucleotides are present in the aligned reads at each locus.&lt;br /&gt;
&lt;br /&gt;
A simple error model is to introduce a small error to the data probability term in the homozygous cases, allowing a small constant probability that nucleotides which don&amp;#039;t match the &amp;#039;&amp;#039;&amp;#039;A&amp;#039;&amp;#039;&amp;#039; allele are observed in the &amp;#039;&amp;#039;&amp;#039;AA&amp;#039;&amp;#039;&amp;#039; case, and respectively a small constant probability that nucleotides not matching the &amp;#039;&amp;#039;&amp;#039;B&amp;#039;&amp;#039;&amp;#039; allele are observed in the &amp;#039;&amp;#039;&amp;#039;BB&amp;#039;&amp;#039;&amp;#039; case. However more sophisticated procedures are available which attempt to more realistically replicate the actual error patterns observed in real data in calculating the conditional data probabilities. For instance, estimations of read quality (measured as [[Phred quality score|Phred]] quality scores) have been incorporated in these calculations, taking into account the expected error rate in each individual read at a locus.&amp;lt;ref name=&amp;quot;Li 2008&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=Mapping short DNA sequencing reads and calling variants using mapping quality scores&lt;br /&gt;
  |author=Li, Heng and Ruan, Jue and Durbin, Richard&lt;br /&gt;
  |journal=Genome research&lt;br /&gt;
  |volume=18&lt;br /&gt;
  |number=11&lt;br /&gt;
  |pages=1851–1858&lt;br /&gt;
  |year=2008&lt;br /&gt;
  |publisher=Cold Spring Harbor Lab&lt;br /&gt;
}}&amp;lt;/ref&amp;gt; Another technique that has successfully been incorporated into error models is base quality recalibration, where separate error rates are calculated - based on prior known information about error patterns - for each possible nucleotide substitution. Research shows that each possible nucleotide substitution is not equally likely to show up as an error in sequencing data, and so base quality recalibration has been applied to improve error probability estimates.&amp;lt;ref name=&amp;quot;Li 2009&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Partitioning of the genotype ====&lt;br /&gt;
&lt;br /&gt;
In the above discussion, it has been assumed that the genotype probabilities at each locus are calculated independently; that is, the entire genotype is partitioned into independent genotypes at each locus, whose probabilities are calculated independently. However due to [[linkage disequilibrium]] the genotypes of nearby loci are in general not independent. As a result, partitioning the overall genotype instead into a sequence of overlapping [[haplotypes]] allows these correlations to be modelled, resulting in more precise probability estimates through the incorporation of population-wide haplotype frequencies in the prior. The use of haplotypes to improve variant detection accuracy has been applied successfully, for instance in the [[1000 Genomes Project]].&amp;lt;ref name=&amp;quot;Abecasis 2010&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=A map of human genome variation from population-scale sequencing.&lt;br /&gt;
  |author=Abecasis, GR and Altshuler, David and Auton, A and Brooks, LD and Durbin, RM and Gibbs, Richard A and Hurles, Matt E and McVean, Gil A and Bentley, DR and Chakravarti, A and others&lt;br /&gt;
  |journal=Nature&lt;br /&gt;
  |volume=467&lt;br /&gt;
  |number=7319&lt;br /&gt;
  |pages=1061–1073&lt;br /&gt;
  |year=2010&lt;br /&gt;
}}&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Heuristic based algorithms ===&lt;br /&gt;
&lt;br /&gt;
As an alternative to probabilistic methods, [[Heuristic (computer science)|heuristic]] methods exist for performing variant calling on NGS data. Instead of modelling the distribution of the observed data and using Bayesian statistics to calculate genotype probabilities, variant calls are made based on a variety of heuristic factors, such as minimum allele counts, read quality cut-offs, bounds on read depth, etc. Although they have been relatively unpopular in practice in comparison to probabilistic methods, in practice due to their use of bounds and cut-offs they can be robust to outlying data that violate the assumptions of probabilistic models.&amp;lt;ref name=&amp;quot;Koboldt 2012&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing&lt;br /&gt;
  |author=Koboldt, Daniel C and Zhang, Qunyuan and Larson, David E and Shen, Dong and McLellan, Michael D and Lin, Ling and Miller, Christopher A and Mardis, Elaine R and Ding, Li and Wilson, Richard K&lt;br /&gt;
  |journal=Genome research&lt;br /&gt;
  |volume=22&lt;br /&gt;
  |number=3&lt;br /&gt;
  |pages=568–576&lt;br /&gt;
  |year=2012&lt;br /&gt;
  |publisher=Cold Spring Harbor Lab&lt;br /&gt;
}}&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Reference genome used for alignment ===&lt;br /&gt;
&lt;br /&gt;
An important part of the design of variant calling methods using NGS data is the DNA sequence used as a reference to align the NGS reads to. In human genetics studies, high quality references are available, from sources such as the HapMap project,&amp;lt;ref name=&amp;quot;Gibbs 2003&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=The international HapMap project&lt;br /&gt;
  |author=Gibbs, Richard A and Belmont, John W and Hardenbol, Paul and Willis, Thomas D and Yu, Fuli and Yang, Huanming and Ch&amp;#039;ang, Lan-Yang and Huang, Wei and Liu, Bin and Shen, Yan and others&lt;br /&gt;
  |journal=Nature&lt;br /&gt;
  |volume=426&lt;br /&gt;
  |number=6968&lt;br /&gt;
  |pages=789–796&lt;br /&gt;
  |year=2003&lt;br /&gt;
  |publisher=Nature Publishing Group&lt;br /&gt;
}}&amp;lt;/ref&amp;gt; which can substantially improve the accuracy of the variant calls made by variant calling algorithms. As a bonus, such references can be a source of prior genotype probabilities for Bayesian based analysis. However in the absence of such a high quality reference, experimentally obtained reads can first be [[sequence assembly|assembled]] in order to create a reference sequence for alignment.&amp;lt;ref name=&amp;quot;Nielsen 2011&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Pre-processing and filtering of results ===&lt;br /&gt;
&lt;br /&gt;
Various methods exist for filtering data in variant calling experiments, in order to remove sources of error/bias. This can involve the removal of suspicious reads before performing alignment and/or filtering of the list of variants returned by the variant calling algorithm.&lt;br /&gt;
&lt;br /&gt;
Depending on the sequencing platform used, various biases may exist within the set of sequenced reads. For instance, strand bias can occur, where there is a highly unequal distribution of forward vs reverse directions in the reads aligned in some neighborhood. Additionally, there may occur an unusually high duplication of some reads (for instance due to bias in [[Polymerase chain reaction|PCR]]). Such biases can result in dubious variant calls - for instance if a read containing a sequencing error at some locus is duplicated due to a PCR bias, that locus will have a high count of the false allele, and may be called as a SNV - and so analysis pipelines frequently filter calls based on these biases.&amp;lt;ref name=&amp;quot;Nielsen 2011&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Methods for detecting somatic variants ==&lt;br /&gt;
&lt;br /&gt;
In addition to methods that align reads from individual sample(s) to a reference genome in order to detect [[germline]] genetic variants, reads from multiple tissue samples within a single individual can be aligned and compared in order to detect somatic variants. These variants correspond to [[genetic mutation|mutations]] that have occurred [[de novo]] within groups of [[somatic cell]]s within an individual (that is, they are not present within the individual&amp;#039;s germline cells). This form of analysis has been frequently applied to the study of [[cancer]], where many studies are designed around investigating the profile of somatic mutations within cancerous tissues. Such investigations have resulted in diagnostic tools that have seen clinical application, and are used to improve scientific understanding of the disease, for instance by the discovery of new cancer-related genes, identification of involved [[gene regulatory networks]] and [[metabolic pathways]], and by informing models of how tumors grow and evolve.&amp;lt;ref name = &amp;quot;Shyr 2013&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=Next generation sequencing in cancer research and clinical application&lt;br /&gt;
  |author=Shyr, Derek and Liu, Qi and others&lt;br /&gt;
  |journal=Biological procedures online&lt;br /&gt;
  |volume=15&lt;br /&gt;
  |number=4&lt;br /&gt;
  |year=2013&lt;br /&gt;
}}&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Recent developments ===&lt;br /&gt;
&lt;br /&gt;
Until recently, software tools for carrying out this form of analysis have been heavily underdeveloped, and were based on the same algorithms used to detect germline variations. Such procedures are not optimized for this task, because they do not adequately model the statistical correlation between the genotypes present in multiple tissue samples from the same individual.&amp;lt;ref name=&amp;quot;Roth 2012&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
More recent investigations have resulted in the development of software tools especially optimized for the detection of somatic mutations from multiple tissue samples. Probabilistic techniques have been developed that pool allele counts from all tissue samples at each locus, and using statistical models for the likelihoods of joint-genotypes for all the tissues, and the distribution of allele counts given the genotype, are able to calculate relatively robust probabilities of somatic mutations at each locus using all available data.&amp;lt;ref name=&amp;quot;Roth 2012&amp;quot;/&amp;gt;&amp;lt;ref name=&amp;quot;Larson 2012&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=SomaticSniper: identification of somatic point mutations in whole genome sequencing data&lt;br /&gt;
  |author=Larson, David E and Harris, Christopher C and Chen, Ken and Koboldt, Daniel C and Abbott, Travis E and Dooling, David J and Ley, Timothy J and Mardis, Elaine R and Wilson, Richard K and Ding, Li&lt;br /&gt;
  |journal=Bioinformatics&lt;br /&gt;
  |volume=28&lt;br /&gt;
  |number=3&lt;br /&gt;
  |pages=311–317&lt;br /&gt;
  |year=2012&lt;br /&gt;
  |publisher=Oxford Univ Press&lt;br /&gt;
}}&amp;lt;/ref&amp;gt; In addition there has recently been some investigation in [[machine learning]] based techniques for performing this analysis.&amp;lt;ref name=&amp;quot;Ding 2012&amp;quot;&amp;gt;{{cite journal&lt;br /&gt;
  |title=Feature-based classifiers for somatic mutation detection in tumour--normal paired sequencing data&lt;br /&gt;
  |author=Ding, Jiarui and Bashashati, Ali and Roth, Andrew and Oloumi, Arusha and Tse, Kane and Zeng, Thomas and Haffari, Gholamreza and Hirst, Martin and Marra, Marco A and Condon, Anne and others&lt;br /&gt;
  |journal=Bioinformatics&lt;br /&gt;
  |volume=28&lt;br /&gt;
  |number=2&lt;br /&gt;
  |pages=167–175&lt;br /&gt;
  |year=2012&lt;br /&gt;
  |publisher=Oxford Univ Press&lt;br /&gt;
}}&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== List of available software ==&lt;br /&gt;
*[http://soap.genomics.org.cn/index.html SOAPsnp]&lt;br /&gt;
*[http://128.32.118.212/thorfinn/realSFS realSFS]&lt;br /&gt;
*[http://samtools.sourceforge.net Samtools]&lt;br /&gt;
*[http://www.broadinstitute.org/gsa/wiki/index.php/The_Genome_Analysis_Toolkit GATK]&lt;br /&gt;
*[http://faculty.washington.edu/browning/beagle/beagle.html Beagle]&lt;br /&gt;
*[http://mathgen.stats.ox.ac.uk/impute/impute_v2.html IMPUTE2]&lt;br /&gt;
*[http://genome.sph.umich.edu/wiki/Thunder MaCH]&lt;br /&gt;
*[http://compbio.bccrc.ca/software/snvmix SNVmix]&lt;br /&gt;
*[http://varscan.sourceforge.net VarScan]&lt;br /&gt;
*[http://gmt.genome.wustl.edu/somatic-sniper Somaticsniper]&lt;br /&gt;
*[http://compbio.bccrc.ca/software/jointsnvmix JointSNVMix]&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&lt;br /&gt;
{{reflist|2}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Category:DNA sequencing]]&lt;/div&gt;</summary>
		<author><name>en&gt;Malcolma</name></author>
	</entry>
</feed>