{"entity": "researcher", "timestamp": "2026-08-20T19:57:20.075Z", "family": "Zhang", "given": "Bo", "initials": "B", "orcid": "0000-0001-8890-8416", "affiliations": ["From the \u2021 Department of Medical Biochemistry and Biophysics, Karolinska Institutet, Scheeles v\u00e4g 2, SE-17177 Solna, Sweden."], "links": {"self": {"href": "https://publications-affiliated.scilifelab.se/researcher/fab1ab9de65a430b97f0a84901798d6e.json"}, "display": {"href": "https://publications-affiliated.scilifelab.se/researcher/fab1ab9de65a430b97f0a84901798d6e"}}, "publications": [{"entity": "publication", "iuid": "b51435ae9ccc4e67baac82324f120ab4", "links": {"self": {"href": "https://publications-affiliated.scilifelab.se/publication/b51435ae9ccc4e67baac82324f120ab4.json"}, "display": {"href": "https://publications-affiliated.scilifelab.se/publication/b51435ae9ccc4e67baac82324f120ab4"}}, "title": "Covariation of Peptide Abundances Accurately Reflects Protein Concentration Differences.", "authors": [{"family": "Zhang", "given": "Bo", "initials": "B", "orcid": "0000-0001-8890-8416", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/fab1ab9de65a430b97f0a84901798d6e.json"}}, {"family": "Pirmoradian", "given": "Mohammad", "initials": "M"}, {"family": "Zubarev", "given": "Roman", "initials": "R", "orcid": "0000-0001-9839-2089", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/5e3f9910c1ff434c8056bdecf537e9ef.json"}}, {"family": "K\u00e4ll", "given": "Lukas", "initials": "L", "orcid": "0000-0001-5689-9797", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/b4464f2bf868498fa6d149a4a6d60e8b.json"}}], "type": "journal article", "published": "2017-05-00", "journal": {"title": "Mol. Cell Proteomics", "issn": "1535-9484", "volume": "16", "issue": "5", "pages": "936-948", "issn-l": "1535-9476"}, "abstract": "Most implementations of mass spectrometry-based proteomics involve enzymatic digestion of proteins, expanding the analysis to multiple proteolytic peptides for each protein. Currently, there is no consensus of how to summarize peptides' abundances to protein concentrations, and such efforts are complicated by the fact that error control normally is applied to the identification process, and do not directly control errors linking peptide abundance measures to protein concentration. Peptides resulting from suboptimal digestion or being partially modified are not representative of the protein concentration. Without a mechanism to remove such unrepresentative peptides, their abundance adversely impacts the estimation of their protein's concentration. Here, we present a relative quantification approach, Diffacto, that applies factor analysis to extract the covariation of peptides' abundances. The method enables a weighted geometrical average summarization and automatic elimination of incoherent peptides. We demonstrate, based on a set of controlled label-free experiments using standard mixtures of proteins, that the covariation structure extracted by the factor analysis accurately reflects protein concentrations. In the 1% peptide-spectrum match-level FDR data set, as many as 11% of the peptides have abundance differences incoherent with the other peptides attributed to the same protein. If not controlled, such contradicting peptide abundance have a severe impact on protein quantifications. When adding the quantities of each protein's three most abundant peptides, we note as many as 14% of the proteins being estimated as having a negative correlation with their actual concentration differences between samples. Diffacto reduced the amount of such obviously incorrectly quantified proteins to 1.6%. Furthermore, by analyzing clinical data sets from two breast cancer studies, our method revealed the persistent proteomic signatures linked to three subtypes of breast cancer. We conclude that Diffacto can facilitate the interpretation and enhance the utility of most types of proteomics data.", "doi": "10.1074/mcp.O117.067728", "pmid": "28302922", "labels": {"Affiliated researcher": null}, "xrefs": [{"db": "pmc", "key": "PMC5417831"}, {"db": "pii", "key": "S1535-9476(20)32397-5"}], "notes": [], "created": "2018-12-05T12:31:06.349Z", "modified": "2026-08-20T09:32:36.111Z"}, {"entity": "publication", "iuid": "b03ed99bfed54a90a9fe776674ff04c6", "links": {"self": {"href": "https://publications-affiliated.scilifelab.se/publication/b03ed99bfed54a90a9fe776674ff04c6.json"}, "display": {"href": "https://publications-affiliated.scilifelab.se/publication/b03ed99bfed54a90a9fe776674ff04c6"}}, "title": "DeMix-Q: Quantification-Centered Data Processing Workflow.", "authors": [{"family": "Zhang", "given": "Bo", "initials": "B", "orcid": "0000-0001-8890-8416", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/fab1ab9de65a430b97f0a84901798d6e.json"}}, {"family": "K\u00e4ll", "given": "Lukas", "initials": "L", "orcid": "0000-0001-5689-9797", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/b4464f2bf868498fa6d149a4a6d60e8b.json"}}, {"family": "Zubarev", "given": "Roman A", "initials": "RA", "orcid": "0000-0001-9839-2089", "researcher": {"href": "https://publications-affiliated.scilifelab.se/researcher/5e3f9910c1ff434c8056bdecf537e9ef.json"}}], "type": "comparative study", "published": "2016-04-00", "journal": {"title": "Mol. Cell Proteomics", "issn": "1535-9484", "volume": "15", "issue": "4", "pages": "1467-1478", "issn-l": "1535-9476"}, "abstract": "For historical reasons, most proteomics workflows focus on MS/MS identification but consider quantification as the end point of a comparative study. The stochastic data-dependent MS/MS acquisition (DDA) gives low reproducibility of peptide identifications from one run to another, which inevitably results in problems with missing values when quantifying the same peptide across a series of label-free experiments. However, the signal from the molecular ion is almost always present among the MS(1)spectra. Contrary to what is frequently claimed, missing values do not have to be an intrinsic problem of DDA approaches that perform quantification at the MS(1)level. The challenge is to perform sound peptide identity propagation across multiple high-resolution LC-MS/MS experiments, from runs with MS/MS-based identifications to runs where such information is absent. Here, we present a new analytical workflow DeMix-Q (https://github.com/userbz/DeMix-Q), which performs such propagation that recovers missing values reliably by using a novel scoring scheme for quality control. Compared with traditional workflows for DDA as well as previous DIA studies, DeMix-Q achieves deeper proteome coverage, fewer missing values, and lower quantification variance on a benchmark dataset. This quantification-centered workflow also enables flexible and robust proteome characterization based on covariation of peptide abundances.", "doi": "10.1074/mcp.O115.055475", "pmid": "26729709", "labels": {"Affiliated researcher": null}, "xrefs": [{"db": "pmc", "key": "PMC4824868"}, {"db": "pii", "key": "S1535-9476(20)33634-3"}], "notes": [], "created": "2018-12-05T09:57:59.735Z", "modified": "2026-08-20T09:32:30.666Z"}]}