Phosphoproteomics workflows in Python

Beta · Python 3.11–3.12

Compare. Enrich. Interpret. Connect.

PhosPy helps you move from phosphosite measurements to biological questions. Compare experimental conditions, explore enriched functions, investigate kinase-associated patterns, and connect individual changes with broader signalling behaviour — while keeping the important analytical choices visible and reproducible.

Install pip install phospy
01

Prepare

Bring your phosphosite measurements and experimental information together in an analysis-ready dataset. PhosPy checks the data along the way so problems can be found before they become biological conclusions.

02

Compare

Ask what changed between your experimental conditions. Test the comparisons that matter to your study, then explore whether the changing phosphosites point towards particular biological functions.

03

Interpret

Follow those changes towards kinase-associated patterns and broader signalling structure. PhosPy helps you move from individual phosphosites towards a more connected view of the response.

Why PhosPy

Focused. Reproducible. Readable.

Phosphoproteomics analysis involves a lot of decisions. PhosPy keeps those decisions visible rather than hiding them behind a pipeline, so you can understand what happened to your data, revisit the assumptions behind an analysis, and repeat the same workflow later.

Scientific questions

Start with the biology.

What changed?

Compare experimental conditions and identify phosphosites whose behaviour differs across your study. The analysis stays centred on the comparisons you designed, whether the experiment is straightforward or paired.

What does it point to?

Put changing phosphosites into biological context with enrichment, then investigate kinase-associated patterns that may help explain the response. The aim is to support interpretation without turning association into certainty.

How does it connect?

Explore how phosphosites group together across the experiment and use signalome analysis to look at broader patterns in the signalling response. The result is a descriptive view of structure, not a claim of causal pathways.

Analysis

For complex studies.

Start with the question

The main workflows follow the questions most phosphoproteomics studies ask: what changed, what is enriched, which kinase-associated patterns stand out, and how do those signals relate to one another?

Handle real experiments

When the study needs more care, PhosPy supports paired designs and specialist approaches for missing data, protein abundance, quantification depth, and unwanted technical variation. These methods stay optional so the analysis only becomes more complex when the experiment calls for it.

Keep the choices visible

PhosPy records the important steps around preprocessing and analysis, including what was kept, changed, or excluded. That makes it easier to understand a result today and to reproduce the same analysis later.

First run

Start with one contrast.

The quickstart uses a small rat dataset to compare treated and control samples. It is a compact way to see how phosphosite data, experimental design, and a differential result fit together before moving to your own study. Once it runs, the workflow guides take you through enrichment, kinase analysis, signalome analysis, and specialist preprocessing when you need it.

import pandas as pd

from phospy import AnalysisReadyDatasetBuilder, DifferentialAnalysisWorkflow
from phospy.advanced import DatasetIntensityTransformConfig
from phospy.api import (
    Contrast,
    DatasetBuildRequest,
    DatasetLocalisationConfig,
    DatasetPreprocessingConfig,
    DifferentialAnalysisRequest,
    ExperimentalDesign,
    Organism,
    SampleDesignRecord,
)

phospho = pd.DataFrame(
    {
        "control_1": [1000.0, 900.0],
        "control_2": [1050.0, 880.0],
        "treated_1": [1800.0, 930.0],
        "treated_2": [1750.0, 920.0],
    },
    index=["MAPK14;Y182;", "GSK3A;S21;"],
)

site_metadata = pd.DataFrame(
    {
        "gene_symbol": ["MAPK14", "GSK3A"],
        "site": ["Y182", "S21"],
        "site_sequence": [
            "LDFGLARHTDDEMTGYVATRWYRAPEIMLNW",
            "PSGGGPGGSGRARTSSFAEPGGGGGGGGGGP",
        ],
        "protein_identifier": ["MAPK14", "GSK3A"],
        "localisation_confidence": [0.95, 0.94],
    },
    index=phospho.index,
)

dataset = AnalysisReadyDatasetBuilder().run(
    DatasetBuildRequest(
        phospho=phospho,
        site_metadata=site_metadata,
        organism=Organism.RAT,
        preprocessing_config=DatasetPreprocessingConfig(
            intensity_transform=DatasetIntensityTransformConfig(policy="log2"),
            localisation=DatasetLocalisationConfig(
                confidence_column="localisation_confidence",
                min_confidence=0.75,
            ),
        ),
    )
)

design = ExperimentalDesign(
    samples=(
        SampleDesignRecord("control_1", "control", "control_r1"),
        SampleDesignRecord("control_2", "control", "control_r2"),
        SampleDesignRecord("treated_1", "treated", "treated_r1"),
        SampleDesignRecord("treated_2", "treated", "treated_r2"),
    )
)

result = DifferentialAnalysisWorkflow().run(
    DifferentialAnalysisRequest(
        dataset=dataset,
        design=design,
        contrasts=(
            Contrast(
                name="treated_vs_control",
                numerator_condition="treated",
                denominator_condition="control",
            ),
        ),
    )
)

print(
    result.table_for("treated_vs_control").loc[
        :, ["display_id", "logFC", "P.Value", "adj.P.Val"]
    ]
)

Input

Phosphosite intensities, information about each measured site, and the sample design for your experiment. Sequence context and localisation evidence help PhosPy check that the data are suitable for the workflows you want to run.

Formats

Work directly with pandas DataFrames or common tabular files such as .csv, .tsv, .txt, and .parquet. PhosPy also provides starting points for MaxQuant and FragPipe/PTMProphet outputs.

Outputs

Each comparison returns its own results table, together with information about what was analysed, what was excluded, and the choices that shaped the result. Later workflows keep the same emphasis on traceable, inspectable outputs.

Interpretation

Know the limits.

PhosPy is designed to help you explore phosphoproteomics data without overstating what the analysis can tell you. Statistical associations remain associations, enrichment depends on the sets and background you provide, kinase results reflect relative support rather than proof of activity, and signalome outputs describe patterns within the experiment rather than causal pathways.

FAQ

Questions, answered.

Is PhosPy a replacement for PhosR?

No. PhosPy is inspired by selected PhosR workflows and brings related ideas into a focused Python package. Its methods and scope are documented independently, so PhosPy should be used for the workflows it supports rather than treated as a one-to-one copy of PhosR.

How do differential analysis and enrichment fit in?

Differential analysis asks what changed between the conditions in your experiment. Enrichment then asks whether the changing sites point towards particular functions or collections. More complex studies can use supported paired designs and specialist modelling options, while enrichment remains an offline analysis using the sets and background you provide.

Which reference data are bundled?

PhosPy 1.7.5 includes an approved rat reference snapshot for supported reference-backed workflows. Human, mouse, and custom analyses need suitable reference data supplied by the researcher, with the source and compatibility information required by the workflow.

What does Signalome need?

Signalome starts from a completed kinase analysis and needs protein-group information for the sites being interpreted. In practice, this means supplying protein_group_id in the site metadata, together with suitable localisation evidence for production analysis.

Where do I go when a first run fails?

Start with the guide for the workflow you are running. The troubleshooting notes walk through the common causes of failure, including data alignment, intensity scale, localisation, missing values, sample design, identifiers, and reference compatibility.

Get started

Bring PhosPy into your next experiment.

Install PhosPy on Python 3.11 or 3.12 and begin with the differential quickstart. From there, follow the workflow that matches the question you want to ask. The documentation keeps the methods, assumptions, examples, and practical guidance close at hand when you need to go deeper.