Peak Fitting (aka Curve Fitting) in Peak® Spectroscopy Software

Infrared, Raman and UV-Vis bands routinely overlap, and the information you want is often carried by sub-bands that never appear as separate maxima. Peak Fitting decomposes a measured band into its individual components and reports the position, height, width and area of each, so they can be interpreted chemically or used for quantitative analysis.

Looking for the theory rather than the software? The Peak Fitting Guide covers where Gaussian and Lorentzian line shapes come from, how to choose between them, what the Levenberg-Marquardt algorithm is actually doing, and worked experiments showing how curve fits fail.

Overview

Peak Fitting uses the Levenberg-Marquardt (LMFit) algorithm, which is widely used for non-linear curve-fitting problems. LMFit is well documented in the literature. From starting estimates it varies peak parameters, calculates a spectrum from those peaks, and evaluates the goodness of fit to the sample spectrum. The metric used to calculate goodness of fit is X2 (Chi-squared), which is the sum of the squared residual spectrum. The residual is the sample spectrum minus the calculated spectrum. LMFit is iterative and requires initial user input. The resulting fit is only as good as the initial peak estimates that are provided. LMFit will always find a solution, but that doesn't mean the solution is optimal or even valid. There can be a number of answers to a non-linear problem. Some solutions may be local minima but are not actually the best fit. This is especially likely if the initial peak estimates are far from the real answer. Another problem is over-fitting. Any spectrum can be fitted provided enough starting peaks, but those peaks may not have a corresponding real peak in the sample. The user is responsible for providing good input and for interpreting the results.

Four peak shapes are available: Gaussian, Lorentzian, Voigt (an even 50:50 blend of the two) and Mixed (the same blend, with the Gaussian fraction under your control). Each peak carries a center, height and full width at half height, and each of those can be left free, bounded to a range, or fixed.

Rules of Thumb

  • Getting the number of peaks right matters more than anything else you do. Every extra peak adds three adjustable parameters and will always improve the numbers, so a low chi-squared can never tell you a peak is real.
  • Place starting centers within about half the spacing between peaks. Inside that range the fit will find them wherever you put them; outside it, a peak can lock onto the wrong band.
  • Don't agonise over starting widths and heights. Determining those is the algorithm's job, and it does it well.
  • With the Mixed shape, fix the Gaussian fraction rather than fitting it. Width and fraction both control the wings of a peak, so the fit trades one against the other and neither ends up determined by your data.
  • Composition percentages are more robust than individual peak parameters. Useful, but it cuts both ways: a believable set of percentages is not evidence that the decomposition behind them is right.
  • A fit can converge, look clean, and still be wrong. Judge it on the residual, and on whether every fitted number is physically sensible.

These are conclusions from experiments on a synthetic spectrum with known component parameters. The experiments and their results are in the Peak Fitting Guide.

Before You Start

  • Peak Fitting works best on small spectral regions containing the overlapping peaks.
  • Before launching the Peak Fitting application, zoom in on the spectral region you want to fit. The displayed region will be used for the peak fitting. Include the whole band plus some flat baseline on either side.
  • If the data is noisy, results may benefit from smoothing before peak fitting. Smooth gently — a heavy smooth changes the very line widths you are trying to measure.
  • If the data has a sloping baseline, the results may benefit from baseline correction before peak fitting. A sloping baseline is the classic cause of phantom peaks at the edges of the region.
  • Fit absorbance rather than transmittance, and ATR correct ATR spectra first.
  • It is important to account for all peaks, including 'buried' peaks.

Tutorial: Protein Secondary Structure

This tutorial is an example of using peak fitting for protein secondary-structure analysis. The broad amide I band (~1700-1600 cm-1) is decomposed into several sub-bands that correspond to alpha-helix, beta-sheet, beta-turn and random-coil motifs. The fitted areas are converted into % secondary structure. It is the classic application of curve fitting in infrared spectroscopy.

  • Load the 'synthetic_protein.spc' datafile into a workspace. It is installed in the 'PeakFit' sub-folder under the PeakSpectroscopy Documents folder.
  • In the 'Analysis' toolbox, choose 'Peak Fitting'.
  • Click the 'Load Peak Fitting Application' button.
  • Click the 'Load Estimates Table' button and select the file 'synthetic_protein.pfit'

'synthetic_protein.spc' was calculated from the seven Lorentzian peaks below, so unlike a real sample the correct answer is known in advance. That makes it a good spectrum to experiment on.

AssignmentCenterAmplitudeFWHH
Beta Sheet1627.90.1019.0
Beta Sheet1636.70.1019.5
Coil1645.70.2520.4
Alpha Helix1656.30.6520.5
turns1667.70.3818.7
turns1678.20.1320.1
Beta Sheet1690.80.0519.5

The Peak Estimates Table

The Estimate Table looks like this. The designation in parentheses after each Center, Height and FWHH value encodes the constraints on that setting. The pair gives the Low Bound Handling first and the High Bound Handling second, using 'U' for Unbounded, 'B' for Bounded and 'F' for Fixed. So (U,U) means the value is free to vary in both directions, and (B,B) means both limits are active.

In this table the Center and FWHH values read (U,U) — the fit may move them anywhere. The Height values read (B,B), constraining every peak height to lie between a Low Bound of 0 and a High Bound of 1. The two limits serve different purposes. The floor at 0 matters: nothing in the mathematics prevents a peak height from going negative during fitting, and a negative component has no physical meaning. The ceiling at 1 simply suits data scaled to around 1 absorbance — raise it if your spectrum has stronger bands, or a tall component will press silently against it. Constraints can be applied using the 'Edit Peak' button.

Initial Peak Estimates for Peak Fitting
synthetic_protein.pfit

Note that these estimates make no attempt to get the heights and widths right — every peak starts at the same width, and the starting heights bear little relation to the real ones. Only the centers are placed deliberately. That is enough.

After loading the spectrum and the estimates table, click the 'Fit Peaks' button. The fitting is performed and the results displayed:

Peak Fit Results
Peak Fit Results.

The Results

From those deliberately rough estimates, the fit recovers the generating parameters exactly — every center, height and width matches the table above to the displayed precision:

AssignmentPeak TypeCenterHeightFWHHArea
Beta SheetLorentzian1627.90000.100019.00002.9845
Beta SheetLorentzian1636.70000.100019.49993.0630
CoilLorentzian1645.70000.250020.40008.0111
Alpha HelixLorentzian1656.30000.650020.500020.9308
turnsLorentzian1667.70000.380018.700011.1621
turnsLorentzian1678.20000.130020.10004.1045
Beta SheetLorentzian1690.80000.050019.50001.5315

Summing the areas by assignment gives the secondary structure composition, which is the point of the exercise:

Structure% of total fitted area
Alpha Helix40.4%
turns29.5%
Coil15.5%
Beta Sheet14.6%

To produce that summary, export the results table to .csv or .xls and total the areas by assignment in a spreadsheet. Naming the peaks in the Edit Peak dialog is what makes this easy — the names travel through to the results table and the exported file, so areas can be grouped without matching wavenumbers by hand. Working in a spreadsheet also leaves you free to exclude a component from the total, which is worth doing when one of the peaks is modeling band shape rather than a distinct species.

This spectrum is noise free and the model has the correct number of components, which is the best case the method can produce. Real data is harder, and it is worth knowing how the same fit behaves when conditions are less ideal. The Peak Fitting Guide works through what happens when peaks are missing, when extra peaks are added, and when the starting positions are displaced.

Manual Peak Selection

The first step in Peak Fitting is to tell the program the nominal positions of the peaks that comprise the spectrum. This can be done manually using the mouse, or by clicking the 'Add Peak' button, or by clicking the 'Find Peaks' button or by loading a saved Peak Estimates table.

Peak Selection with the Mouse

It helps to overlay the 2nd derivative by checking the '2nd derivative' box. The 2nd derivative is useful for finding buried peaks. A minimum in the 2nd derivative is the location of a peak.

Spectrum with the second derivative overlaid, showing minima at the positions of buried peaks
Spectrum with overlaid 2nd Derivative

To manually create a peak, right-click the mouse at a peak location. A peak marker is created at the mouse location. The peak marker consists of three elements: the Center marker and two width markers. The width initially is set to the default width provided in the Peak Options table.

creating peak markers for peak fitting
A Peak Marker.

The position, height, and width of the peak can be changed by dragging a marker with the left mouse button. When positioned over the Center Marker, the mouse cursor changes to indicate that the peak can be moved up and down and left and right. Moving the Center Marker moves the Width Markers along with it.

The peak center marker
The Peak Center Marker.

When positioned over a vertical width marker, the mouse cursor changes to indicate that the width marker can be moved. Moving a width marker automatically moves the other width marker as well but leaves the Center Marker where it is.

The peak width markers
The Peak Width Markers.

When a peak marker is created with the mouse, an entry in the 'Peak Estimates' table is made corresponding to that peak. To remove a peak marker using the mouse, position the mouse over either the center marker or one of the width markers, and right-click. Also, the row corresponding to the peak in the Peak Estimates table can be selected and then the 'Remove Peak' button clicked. The left mouse button can still be used to expand (zoom in) on the spectral display, as long as it is not on a marker when the left mouse button is pressed.

Adding Peaks with the Keyboard

Peaks can also be added manually by clicking the 'Add Peak' button. This dialog box appears:

Editing Peak Properties
Editing a Peak.

The peak Center, Height, and FWHH can be entered manually. In addition, a peak can be given a name. A name can be useful in the analysis of the peaks. For instance, in this synthetic protein spectrum, peaks can be assigned to Beta Sheets, Alpha Helices, Turns, and so forth. The 'Bound Handling' entries allow for restricting the peak search. The choices are:

UnboundedAllow the 'Value' to vary without any restrictions.
BoundedRestrict the 'Value' to be in the range of 'Low Bound' and/or 'High Bound'.
FixedDo not allow the 'Value' to change during the peak fitting optimization.

The 'Low Bound' value is only applied when 'Low Bound Handling' is set to 'Bounded' and the 'High Bound' value is only applied when 'High Bound Handling' is set to 'Bounded'. Note: there is nothing to restrict a peak height from becoming negative during LMFit. So, by default the Height Low Bound is set to 0, and the Height Low Bound Handling is set to Bounded.

Bounds are a claim about the answer, not a hint. An unbounded parameter is a suggestion the data is free to overrule. A bounded one cannot leave the range you set, so the range needs to come from independent knowledge rather than from your guess. A fixed one the fit cannot contradict at all. If a fitted value comes to rest exactly on one of its limits, the optimizer wanted to keep going and was not allowed to — treat that number as a property of your settings, not of the sample.

Automatic Peak Selection

Clicking the 'Find Peaks' button will perform a 2nd derivative analysis of the spectrum and select peaks on that basis. 'Find Peaks' can be useful, but manually selecting peaks usually yields better results because the eye of an analyst is better at discerning fine structure than a computer algorithm. In this graphic, the 2nd derivative overlay was used to select the positions of the peaks:

Automatic Initial Peak Finding for Peak Fit
Automatic Peak Finding.

And the table of Peak Estimates looks like this:

The table of peak estimates from automatic peak finding
Automatic Peak Finding.

Treat automatic peak finding as a starting point to be edited rather than as an answer. It places a peak at every minimum of the second derivative, which usually means more peaks than the chemistry justifies, and on noisy data some of those minima are noise. Delete the ones you cannot justify before the first fit, and add them back only if the residual asks for them.

Automating Peak Fitting with Batch Processing

The batch processing that is built into Peak can be used to automate peak fitting. This allows unattended analysis of samples, and the ability to analyze many samples at once. To set this up,

  • First, create a peak estimate table as described above and save a .pfit peak fitting model.
  • Click on 'Batch Processor' in the 'Advanced' toolbox.
  • Click on 'Create a new Sequence'
  • Click on the 'Analysis' tab.
  • Highlight 'Peak Fitting' and click the 'Add to Sequence' button. The peak fitting options table appears.
  • Load a peak estimate table .pfit file.
  • To save the calculated peaks as individual .csv files, click the 'Save Peaks' checkbox, and choose a folder as the destination for these files. The files will be saved in a sub-folder (named after the sample filename) of the destination folder. The table of results is also saved.
  • Save the batch processing sequence.

If smoothing or baseline correction need to be done prior to the peak fitting, these manipulations can be added to the sequence.

Batch processing is worth setting up for any study with more than a handful of samples, and not only to save time. Applying one saved model, with one set of constraints, to every spectrum removes the sample-to-sample operator variation that makes fitted results hard to compare. See Batch Processing for the general mechanism, and Python Scripting if you need to post-process the exported peak tables.

Typical Applications

  • Protein secondary structure by FTIR — decomposing the amide I envelope into alpha-helix, beta-sheet, beta-turn and random-coil sub-bands.
  • Polymer crystallinity — resolving overlapping crystalline and amorphous bands.
  • Carbon materials by Raman — separating the overlapping D and G bands.
  • Silicates, glasses and minerals — decomposing broad Si-O stretching envelopes.
  • Cellulose, lignin and biomass — separating heavily overlapped C-O and O-H regions.
  • Quantitative analysis of overlapped bands — recovering the area of an analyte band that sits partly under an interferent.

Curve fitting is complementary to the multivariate methods in Peak. Peak fitting is the right tool when you want physically interpretable parameters for individual bands; Classical Least Squares and Partial Least Squares are the right tools when you want the most robust prediction of concentration from whole spectra. Spectral subtraction is often a simpler answer when a single known interferent is the only problem.

Frequently Asked Questions

What is peak fitting in spectroscopy?
Peak fitting, also called curve fitting or band deconvolution, represents a measured spectral band as the sum of several mathematical peak functions. An optimization algorithm adjusts the position, height and width of each component until the sum reproduces the measured spectrum, and the resulting parameters are then interpreted chemically.
How many peaks should I use?
The smallest number that leaves a residual resembling instrument noise. This is the decision that matters most, and no goodness-of-fit number will make it for you.
How accurate do my starting estimates need to be?
Heights and widths barely matter. Positions matter more, but the tolerance is wide — roughly half the spacing between adjacent peaks. See the experiments for the measurements behind that.
Which peak shape should I choose?
Gaussian for solids, powders, polymers and glasses; Lorentzian for low pressure gases and many mobile liquids; Voigt for most condensed phases. If unsure, fit an isolated band in the same spectrum with each shape and keep the one that leaves the flattest residual. The Peak Fitting Guide explains why.
Height or area for quantitation?
Area, in most cases. Broadening reduces height without changing area, so area-based calibrations are usually more linear.
Is curve fitting the same as deconvolution?
No. Curve fitting leaves the measured data untouched and builds a model to match it. Deconvolution alters the data itself by attempting to remove instrumental broadening.
How do I know the fit is good?
Judge the residual against the noise in a flat region of the same spectrum, and check that every fitted parameter is physically sensible. A low chi-squared on its own proves nothing.
Why do two analysts get different answers from the same spectrum?
Because the number of peaks, the line shape and the constraints are all user choices. This is why saving a .pfit model and applying it by batch processing is so much more reproducible than fitting each spectrum by hand.

Related Pages