ClaridaClarida

Why do you need to use the real PCR efficiency in the analysis of your qPCR data?

All posts · · J.M. Ruijter & Maurice van den Hoff

Schematic illustrations of PCR show that in each amplification cycle the number of target molecules is doubled. In such an ideal world, we would not need to talk about the PCR efficiency, as it would be 2 or 100% in all reactions. However, in the real world this is not the case. At best, the amplification efficiencies of reactions amplifying the same target are the same, but different targets often have different amplification efficiencies. This difference is the result of differences in primer sequence and 3′-terminal stability, secondary structure in the template or amplicon, genetic variation between specimens, and variation in polymerase kinetics due to buffer composition and inhibitors in samples from different clinical or environmental conditions. The PCR efficiency is not always, in fact hardly ever, the optimal 100%, yet it is still commonly ignored. For this reason, the recent MIQE update strongly recommends that qPCR data analyses include a so-called 'efficiency-correction' approach. In this blog, we will show the effect of ignoring the PCR efficiency and show that by NOT taking the PCR efficiency of the assays into account, one will never obtain the correct results from a qPCR study.

Cq depends on the PCR efficiency

Shortly back to the qPCR basics that we described in our first blog. The exponential phase of the PCR reaction is determined by the kinetic equation of PCR: Nc=N0⋅EcN_c = N_0 \cdot E^{c}. The common way to analyse the observed fluorescence data is that, after baseline correction, the user, or the PCR instrument software, sets a quantification threshold at a chosen fluorescence value and determines where this threshold intersects each of the amplification curves. The line from this intersection down to the X-axis intersects this cycle axis at the so-called quantification cycle or Cq value (still often called the threshold cycle, or Ct). The Cq value is thus the fractional number of cycles that the reaction needed to reach the fluorescence threshold. Unfortunately, in most clinical and many experimental research reports, this Cq value is reported as the outcome of a PCR despite the fact that this Cq value is dependent on the PCR efficiency, as we showed in the first blog (Fig. 1).

Figure 1. Illustration of the dependence of Cq on the PCR efficiency. For two reactions with the same input (N0=30N_0 = 30) but a different amplification efficiency (orange E=1.9E = 1.9, blue E=1.8E = 1.8), the reported Cq values differ. With this difference in Cq values (dCq=−2.81dCq = -2.81), assuming an efficiency of 100% for both reactions, would suggest a 7-times (=2−dCq=22.81= 2^{-dCq} = 2^{2.81}) difference in abundance level.

From F0 to the expression ratio

To include the PCR efficiency in this analysis per reaction, the fluorescence at the start of the reaction (F0F_0) can be calculated with the threshold (FqF_q), Cq, the efficiency value (EE) and the inverse of the kinetic PCR equation (F0=Fq/ECqF_0 = F_q / E^{Cq}). This F0F_0 is in fact the primary outcome of the PCR reaction and, because of the linear relation between number of targets and their associated fluorescence, F0F_0 is a direct measure of N0N_0, the number of targets at the start of the reaction.

However, for most of you, F0F_0 and N0N_0 will be new concepts because you are used to see the outcome of a PCR analysis reported as an expression ratio between target-of-interest and reference genes, calculated as 2−dCq2^{-dCq}, or maybe even as a fold difference, or treatment effect, between an experimental group and a control group, calculated as 2−ddCq2^{-ddCq}. The latter is the comparative Ct method, or Livak method, better known as 2−ΔΔCt2^{-\Delta\Delta Ct} (Livak and Schmittgen, 2001). How do we go from the above defined F0F_0 to these dCq equations?

In the history of qPCR data analysis, we can distinguish different simplification steps that were taken in the analysis; here we will discuss these 'solutions' with respect to their assumptions, problems and caveats. The first problem is the variation between different clinical or biological samples due to differences in sample size and yield of pre-PCR processing of the samples. Assuming that sampling and pre-processing affect every target in a sample in the same way, qPCR results are preferably reported as an expression ratio (ER) between target-of-interest (toi) and reference (ref) genes [Eq. 1]. With the kinetic equation for F0F_0 we can write this equation for ER as [Eq. 2]. In the dCq approach, the equation for ER is simplified to [Eq. 3].

ER=F0,toiF0,refER = \frac{F_{0,toi}}{F_{0,ref}}[Eq. 1]
ER=Fq/Etoi CqtoiFq/Eref CqrefER = \frac{F_q / E_{toi}^{\,Cq_{toi}}}{F_q / E_{ref}^{\,Cq_{ref}}}[Eq. 2]
ER=2 (Cqref−Cqtoi)=2−dCqER = 2^{\,(Cq_{ref} - Cq_{toi})} = 2^{-dCq}[Eq. 3]

Two assumptions hidden in the simplification

Note that this simplification assumes two things: first, that the quantification threshold (FqF_q) is the same for both targets and, second, that both assays amplify with 100% efficiency: both EtoiE_{toi} and ErefE_{ref} are replaced by 2. The validity of the reported ER depends on the validity of both assumptions. In this blog, we show equations with only 1 reference but in good PCR practice (according to the MIQE guidelines) you need a stable set of at least two validated reference genes.

The second simplification that entered the field of qPCR analysis involved the practice to mainly report the fold-difference between the gene expression ratios in a control group and a group that underwent an experimental condition. This was triggered by the use of qPCR in experimental research studies about the difference in gene expression in various experimental groups.

Starting with the above expression ratio (ER), such a fold-difference (FD) between groups can be written as [Eq. 4].

FD=2−dCqtreated2−dCqcontrol=2−ddCqFD = \frac{2^{-dCq_{treated}}}{2^{-dCq_{control}}} = 2^{-ddCq}[Eq. 4]

The Pfaffl efficiency correction

When we would have started this derivation of FD with the original F0F_0 equation, you would have seen that there is no longer the prerequisite that Fq,toiF_{q,toi} and Fq,refF_{q,ref} are set at the same fluorescence level. This cancelling out of the threshold setting from the equations can be considered a bonus of reporting a fold difference.

Soon after the publication of the 2−ddCq2^{-ddCq} approach, a paper by Pfaffl (2001) recognised that in qPCR practice the assumption that all targets were amplified with 100% efficiency did not hold. This paper, therefore, recommended an efficiency-corrected approach based on the difference of Cq values between target-of-interest and reference genes in control and experimental groups [Eq. 5].

FD=Etoi (Cqtoi,control−Cqtoi,treated)Eref (Cqref,control−Cqref,treated)FD = \frac{E_{toi}^{\,(Cq_{toi,control} - Cq_{toi,treated})}}{E_{ref}^{\,(Cq_{ref,control} - Cq_{ref,treated})}}[Eq. 5]

What the simplifications cost

Note that in this equation, the difference in Cq values is not between target and reference, as in the 2−ddCq2^{-ddCq} equation, but between treated and control groups per assay. Although the 'Pfaffl' equation thus does not ignore differences in amplification efficiency, it uses mean Cq values for each target in each group. This averaging of Cq values leads to loss of information on failed reactions and information on biological variation within the experimental groups.

The above description of the main simplification steps in qPCR analysis shows that fundamentally the 2−dCq2^{-dCq} and 2−ddCq2^{-ddCq} approaches are based on the basic equation for PCR kinetics and its inverse equation to calculate the F0F_0 per reaction: F0=Fq/ECqF_0 = F_q / E^{Cq}. However, for convenience and simplification, these approaches assume the amplification efficiency, EE, to be 2, or 100%, in all reactions for all targets in every experimental condition. To further muddle the issue, in many papers the '2 to the power minus' part of these equations is omitted as a further simplification and just ddCq values are published as if they are a fold-difference. Such papers leave it to the reader to figure out whether the Y-axis annotation saying 'fold difference' is indeed showing fold difference values or just the ddCq value. In case of the latter, the reader has to assume a PCR efficiency of 2, decide whether the axis shows ddCq or −ddCq, and do the math. It is of the utmost importance to realize that even when the efficiency values are mentioned in the paper, as required by the original MIQE guidelines, these values cannot be applied because the reported ddCq values result from two different targets. To do so, the individual Cq values are required.

A simulation: two targets, two conditions

To illustrate the effect of ignoring the PCR efficiency in qPCR data analysis, we performed a simulation with two targets and two experimental conditions. In this simulation, the PCR efficiencies of the target-of-interest (toi) and the reference target (ref) both varied from 1.7 to 2.0 (or from 70% to 100%). In the control condition, the number of starting copies of the toi and ref are set to 50 and 250 copies, respectively (Fig. 2A, left). In the treated condition, due to a sampling mistake introduced in the simulation, the sample input is 5-times larger, and the PCR of the reference target, by definition not affected by the treatment, thus starts with 1250 copies. Therefore, because of the treatment effect of 25 times, the toi starts at 6250 copies (Fig. 2A, right). To also illustrate the effect of setting FqF_q at different levels, we simulated thresholds at three fluorescence levels corresponding to 10910^{9}, 101010^{10} and 101110^{11} copies of the amplicons (Fig. 2A).

Figure 2A. Simulated amplification curves in control (left) and treated (right) conditions. The starting numbers of the target of interest (toi, blue) and of the reference (ref, orange) in the control condition are 50 and 250, respectively, resulting in a toi/ref ratio of 0.2. The treatment effect is 25 times and due to sampling error, the treated sample is 5 times larger than the control sample. Combined these effects result in starting numbers of 6250 and 1250 for toi and ref, respectively, in the treated condition. The amplification curves for PCR efficiencies of 1.7 (dotted) and 2.0 (solid) for both assays are shown. It is of relevance to realize that the amplification curves for PCR efficiencies between 1.7 and 2.0 are found in between the solid and dotted lines. The quantification thresholds (green lines) were set at 10910^{9}, 101010^{10} or 101110^{11} copies of the amplicon.

Expression ratios: a 20,000-fold range

Using those quantification thresholds, we determined all Cq values and calculated the dCq values between toi and ref for every combination of efficiency values in the control and treated condition (Fig. 2B). In these graphs we see that the dCq values depend on the threshold level and that even when the PCR efficiencies are the same, the dCq values are different rather than constant. Each dCq value was used in the 2−dCq2^{-dCq} equation to calculate the toi/ref expression ratio in the control and treated condition (Fig. 2C). Remember that in this simulation for the control condition this ratio was set at 0.2 (= 50/250) and in the treated condition this ratio is 5 (= 6250/1250). Note that only when both PCR efficiencies are 2, the calculated ratios are indeed 0.2 and 5, respectively. For all other efficiency combinations, the expression ratio calculated with 2−dCq2^{-dCq} is not correct and results in a 20,000-fold range of ER values in both conditions. Note that the reported expression ratio can be in the wrong direction, suggesting over-expression rather than the true under-expression of the target in the control condition, or vice versa in the treated condition.

Figure 2B. dCq values in the control and the treated condition. The graphs show all dCq values that will be observed between the toi and the ref in the control and treated samples with the given range of efficiencies (toi on X-axis, ref as green lines) and quantification thresholds (groups of lines) in Figure 2A.
Figure 2C. Gene expression ratios in the control and the treated condition. All gene expression ratios were calculated from the dCq values in Figure 2B using the 2−dCq2^{-dCq} equation. Note that in the control condition (left panel) the true expression ratio is 0.2 but the calculated ratios vary between 0.005 and 100. Similarly, the expected ratio of 5 (0.2 times the treatment effect of 25) in the treated condition (right panel) ranges from 0.05 to 1000.

Fold differences: wrong unless both efficiencies are 2

The dCq values (Fig. 2B) were then used to calculate the ddCq value for every combination of efficiency values (Fig. 2D, left) and these ddCq values were used in the 2−ddCq2^{-ddCq} equation to calculate the fold difference between treated and control conditions. Remember that the treatment effect set in the simulation was 25 times (Fig. 2A). This expected fold difference value is only found when both the target-of-interest and the reference are amplified with 100% or 2.0 efficiency. For all other efficiency values, the result is wrong: the lower the actual efficiency, the larger this error is. Because the error in the expression ratio in treated and control conditions is correlated (visible as the similarity of the graphs in the left and right panel of Fig. 2C), the error in fold difference is less than the error in expression ratio (Fig. 2C) but still ranges between 40% too low to over 4 times too high. Because the actual error is dependent on both the actual efficiency values and gene expression levels, the direction of the error in fold difference cannot be inferred from the actual efficiency values.

Figure 2D. ddCq values and fold differences between treated and control condition. The left panel shows the observed ddCq values between treatment and control conditions calculated from the dCq values in Figure 2B. As mentioned in the text, in the derivation of the 2−ddCq2^{-ddCq} equation the quantification threshold cancels out. Consequently, there is only one ddCq value per efficiency combination. The right panel shows the fold difference, or treatment effect, calculated from these ddCq values as 2−ddCq2^{-ddCq}. Remember that the given treatment effect is 25 times (Figure 2A) which is only observed when both efficiency values are 2.0.

Assuming 100% efficiency always gives a wrong result

Taken together, the simulation and graphs in Figure 2 show that assuming that the PCR efficiency is 2, or 100%, for all assays while that is not the case, will always give a wrong result for the toi/ref expression ratio as well as for the fold difference, or treatment effect, between control and treated conditions. Note that similar wrong results would occur when the simulation was carried out with a set of reference genes, each with their own amplification efficiency. Because these errors compound exponentially, what begins as a seemingly small methodological assumption becomes an analytical distortion. The simulation shows that in some cases this can even lead to a conjectured treatment effect that is in the reverse direction of the true effect when no efficiency correction is applied. Although it is often stated that these errors are acceptable when PCR efficiency differences are small, there is no reason at all to accept these errors. With the same data and similar effort, the analysis can be carried out with the true PCR efficiencies and reveal the correct result, as will be shown below.

Efficiency-corrected analysis with F0

Using the inverse of the kinetics equation of PCR, the fluorescence associated with the number of target copies at the start of the reaction, F0F_0, can be calculated easily from the Cq value per reaction and the PCR efficiency of the assay (see a future blog on how to determine the PCR efficiency). Figure 3, based on true qPCR data, shows how this analysis is performed. The experiment consists of two experimental groups with four biological replicate samples in each group. There are no technical PCR replicates. The amplification curves of these reactions are shown (Fig. 3, graphs). Starting with the F0F_0 values for each target per reaction, the toi/ref ratio per sample, the mean ratio per group and the treated/control fold difference between groups (Fig. 3, table) can be calculated. Note that the toi/ref ratios per group can be compared with a Mann-Whitney test to determine whether the expression ratio in the treated group has been significantly affected by the experiment.

Although this test already answers the biological or clinical question, it is common practice to report a fold difference or treatment effect. This fold difference (FD) can be calculated by dividing the geometric mean toi/ref ratios in the treated and control group (Fig. 3, table). Error propagation can then be used to calculate an SD of this fold difference from the SD and the mean expression ratio (ER‾\overline{ER}) of each group [Eq. 6].

(SDFDFD)2=(SDtreatedER‾treated)2+(SDcontrolER‾control)2\left( \frac{SD_{FD}}{FD} \right)^{2} = \left( \frac{SD_{treated}}{\overline{ER}_{treated}} \right)^{2} + \left( \frac{SD_{control}}{\overline{ER}_{control}} \right)^{2}[Eq. 6]

Per-sample analysis reveals deviating samples

Note that in this example the SD of the fold difference is rather large due to one very high toi/ref ratio in the treated group (Fig. 3, yellow cell in the table). The ability to identify this deviating sample shows an additional advantage of starting the analysis with the F0F_0 per sample, compared to the ddCq or Pfaffl's efficiency-corrected approach. In the latter approaches, the mean Cq value per target and group is used which leads to loss of information on individual samples and loss of variation within the groups. With the illustrated F0F_0 approach, this loss of information does not happen and a quality control and outlier detection at every level of the experiment can be carried out.

Figure 3. Effect of quantification method on fold-change estimation from the same qPCR data. Left: Amplification curves for target of interest (toi; orange) and reference (ref; blue) assays in four control samples are shown. Right: Amplification curves for the same assays in four treated samples are shown. The PCR efficiency values are 1.716 or 72% for toi and 1.907 or 91% for ref. The table shows the experimental design (group, sample and target, columns 1–3), Cq values and F0F_0 values (columns 4–5). The F0F_0 values per sample are used to calculate the toi/ref ratio (column 6). The non-parametric two-sample Mann-Whitney test was performed on these toi/ref ratios of the treated and control groups and shows a significant difference in the expression ratios per group. The geometric mean toi/ref ratio per group (column 7) and their standard deviation (column 8) can then be used to calculate the fold difference between groups. Data are a selection from the CLSTN1 (target) and HPRT1 (reference) assays in the Vermeulen et al. dataset (Lancet Oncol 2009).

Use the real PCR efficiency

We hope that this blog convincingly showed you that when you do NOT use the real PCR efficiency values in your qPCR analysis, the results will suffer from an error of unknown magnitude and may even be in the wrong direction. As a consequence, the results that you are reporting are meaningless. Efficiency-correction is the only way to avoid these errors and to reach correct and reliable results. So to report valid and meaningful results from your experiment in the future, you have to use the Cq and efficiency values that you obtain from your qPCR analysis and apply these in the efficiency-corrected analysis approach that we described here. Because this efficiency-corrected analysis is initially performed per sample, the added bonus of this approach is that you will be able to detect deviating samples before losing this information by calculating mean expression levels per group.

Key takeaways

  • The PCR efficiency is hardly ever 100%, and different assays often amplify with different efficiencies.
  • The dCq and ddCq methods assume 100% efficiency for every assay. When that is not true, expression ratios and fold differences are wrong by an unknown amount and can even point in the wrong direction.
  • Published ddCq values cannot be corrected afterwards, even when the efficiencies are reported. That requires the individual Cq values.
  • Calculate F0 per reaction from its Cq value and the real efficiency of the assay, then build expression ratios and fold differences from these F0 values.
  • Analyzing per sample keeps the biological variation and lets you spot deviating samples before averaging per group.

J.M. Ruijter

Retired Principal Investigator, AMC Amsterdam

Developed methods for 3D analysis and visualization of gene expression patterns during embryogenesis and for the analysis of quantitative PCR data. Creator of LinRegPCR, a widely used method for assumption-free estimation of PCR amplification efficiency.

Maurice van den Hoff

Retired Associate Professor, Amsterdam UMC

Led a research group focused on the developmental mechanisms of normal and abnormal cardiac development and the molecular response of the diseased heart, in particular after cardiac infarction. With Jan Ruijter, he has worked to improve the reliability of quantitative PCR data analysis.

Analyze with your real efficiencies.

Clarida determines the PCR efficiency of every assay and uses it in your quantification, sample by sample. Try it on your own data. Free to start, no install.