MetaX Cookbook

This guidebook is for the MetaX GUI version. If you are using the CLI, we recommend reading the documentation for instructions on how to use each MetaX module from the command line.

Overview

MetaX is a novel tool for linking peptide sequences with taxonomic and functional information in Metaproteomics. We introduce the Operational Taxon-Function (OTF) concept to explore microbial roles and interactions ("who is doing what and how") within ecosystems.

MetaX also features statistical modules and plotting tools for analyzing peptides, taxa, functions, proteins, and taxon-function contributions across groups, and can now export recorded GUI analysis steps as runnable workflow notebooks for reproducible downstream use.

abstract

Project Page

Visit GitHub to get more information:

https://github.com/byemaxx/MetaX

Getting Started

main_window tools_menu


Exploring Data with MetaX

See the Preparing Your Data section to build the database and annotate peptides to OTFs before starting.

Module 1. OTF Analyzer

After obtaining the Operational Taxa-Functions (OTF) Table using the Peptide Annotator, you can perform downstream analysis with the OTF Analyzer.

1. Data Preparation

OTFs (Operational Taxa-Functions) Table: Obtained from the Peptide Annotator module.

Meta Table: The first column is sample names, and the other columns represent different groups. If no meta table is provided, meta info will be generated automatically: (1) all samples are in the same group; (2) each sample is a separate group.

Example Meta Table:

samples Individuals Treatment Sweetener
sample_1 V1 Treatment XYL
sample_2 V1 Treatment XYL
sample_3 V1 Treatment XYL
sample_4 V1 Control PBS
sample_5 V1 Control PBS
sample_6 V1 Control PBS

You can load example data by clicking the button.

load_example

Then, click Go to start the analysis.

2. Data Overview

The Data Overview provides basic information about your data, such as the number of taxa, functions, and proportions.

data_overview data_overview_func data_overview_filter

3. Set TaxaFunc

set_multi_table

Data Selection

FUNC_prop

Sum Proteins Intensity

Click Generate Protein Intensity Table to sum peptides to proteins if the Protein column is in the original table.

Data preprocessing

There are several methods for detecting and handling outliers.

In all methods, you can choose one meta column for outlier detection and another meta column for handling outliers.

You can choose outlier imputation by each group or by all samples.

If you use Z-Score, Mean centring, or Pareto Scaling for data normalization, the data will be given a minimum offset again to avoid negative values.

Then, click Go to create a TaxaFunc object for analysis.

TaxaFunc_ready

Then you can check the tables in the Table Review section and export them.

table_review table_review_open_window

4. Basic Stats

PCA, Correlation and Box Plot

basic_stats_pca

You can select meta groups or samples (default: all) to plot PCA, Correlation, and Box Plot for Taxa, Function, Taxa-Func, Peptide, and Protein tables.

pca pca_3d correlation boxplot

Heatmap and Bar Plot

add_to_list add_top_list add_a_list heatmap_original basic_stats_bar basic_stats_bar_setting

Peptide Query

peptide_query

5. Cross Test

T-TEST

t_test

ANOVA-TEST

anova_test

Significant Taxa-Func

Plot Cross Heatmap

t_test_res corss_heatmap_setting corss_heatmap t_test_heatmap

Group-Control TEST

Set a Group as "Control", then compare all groups to Control

Bingo! You noticed the hidden function of MetaX, click Help -> About -> Like 3 times to unlock the function to compare all groups to control.

Differential Expression (Limma / DESeq2)

(Ultra-Up(Down): |log2FC| > Max log2FC)

Tukey Test

tukey_test taxa_func_linked_only tukey_plot

6. Expression Analysis

Co-Expression Networks & Heatmap

image-20230728142905839 image-20230728143058568 co_network_pic image-20230728152236517 image-20230728150853953 bar_switch_satck bar_to_line

Taxa-Func Network

taxa_func_network

8. Restore Last TaxaFunc Object

Preparing Your Data

Module 2. Database Builder

Note: The results from MetaLab v2.3 MaxQuant workflow do not require database building. However, we do not recommend using these results as input to MetaX, as many peptides may be discarded.

Option 1: Build Database Using MGnify Data

Ensure you download the correct database type corresponding to your data.

MetaX supports the MGnify catalogues listed in the Database Builder selector, including barley-rhizosphere, human-skin, maize-rhizosphere, marine-sediment, soil, and tomato-rhizosphere. The selector and command-line options are generated from MetaX's supported-source list. marine-eukaryotes is intentionally not enabled by default because it is a beta eukaryotic catalogue with an eggNOG annotation caveat.

dbbuilder

Option 2: Build Database Using Own Data

  1. Annotation Table: A TSV table (tab-separated), with the first column as protein name joined with Genome by "_", e.g., "Genome1_protein1", and other columns containing annotation information.
dbbuilder_own
  1. Taxa Table: A TSV table (tab-separated), with the first column as Genome name, e.g., "Genome1", and the second column as taxa.

Example Annotation Table:

Query Preferred_name EC KEGG_ko
MGYG000000001_00696 mfd - ko:K03723
MGYG000000001_02838 hxlR - -
MGYG000000001_01674 ispG 1.17.7.1,1.17.7.3 ko:K03526
MGYG000000001_02710 glsA 3.5.1.2 ko:K01425
MGYG000000001_01356 mutS2 - ko:K07456
MGYG000000001_02630 - - -
MGYG000000001_02418 ackA 2.7.2.1 ko:K00925
MGYG000000001_00728 atpA 3.6.3.14 ko:K02111
MGYG000000001_00695 pth 3.1.1.29 ko:K01056
MGYG000000001_02907 - - ko:K03086
MGYG000000001_02592 rplC - ko:K02906
MGYG000000001_00137 - - ko:K03480,ko:K03488

Example Taxa Table:

Genome Lineage
MGYG000000001 d_Bacteria;p_Firmicutes_A;c_Clostridia;o_Peptostreptococcales;f_Peptostreptococcaceae;g_GCA-900066495;s_GCA-900066495 sp902362365
MGYG000000002 d_Bacteria;p_Firmicutes_A;c_Clostridia;o_Lachnospirales;f_Lachnospiraceae;g_Blautia_A;s_Blautia_A faecis
MGYG000000003 d_Bacteria;p_Bacteroidota;c_Bacteroidia;o_Bacteroidales;f_Rikenellaceae;g_Alistipes;s_Alistipes shahii
MGYG000000004 d_Bacteria;p_Firmicutes_A;c_Clostridia;o_Oscillospirales;f_Ruminococcaceae;g_Anaerotruncus;s_Anaerotruncus colihominis
MGYG000000005 d_Bacteria;p_Firmicutes_A;c_Clostridia;o_Peptostreptococcales;f_Peptostreptococcaceae;g_Terrisporobacter;s_Terrisporobacter glycolicus_A
MGYG000000006 d_Bacteria;p_Firmicutes;c_Bacilli;o_Staphylococcales;f_Staphylococcaceae;g_Staphylococcus;s_Staphylococcus xylosus
MGYG000000007 d_Bacteria;p_Firmicutes;c_Bacilli;o_Lactobacillales;f_Lactobacillaceae;g_Lactobacillus;s_Lactobacillus intestinalis
MGYG000000008 d_Bacteria;p_Firmicutes;c_Bacilli;o_Lactobacillales;f_Lactobacillaceae;g_Lactobacillus;s_Lactobacillus johnsonii
MGYG000000009 d_Bacteria;p_Firmicutes;c_Bacilli;o_Lactobacillales;f_Lactobacillaceae;g_Ligilactobacillus;s_Ligilactobacillus murinus

Module 3. Database Updater

The Database Updater allows updating the database built by the Database Builder or adding more annotations. This step is optional.

db_updater

Option 1: Built-in Mode

Built-in dbCAN_seq mode merges precomputed annotations by exact protein ID; it does not run sequence-similarity searches or re-annotate custom proteins. Incoming annotation columns replace existing columns with the same names, and MetaX writes a warning listing the replaced columns. For a custom protein database, run dbCAN/run_dbCAN on your own protein FASTA and import the resulting TSV with matching MetaX protein IDs using Option 2. Built-in sources are available from dbCAN_seq.

Option 2: TSV Table

Extend the database by adding a new database to the database table. Ensure the column separator is a tab and the first column is the Protein name, with other columns containing function annotations.

Example:

Protein ID COG KEGG ...
MGYG000000001_02630 Function 1 Function 1 ...
MGYG000000001_01475 Function 2 Function 1 ...
MGYG000000001_01539 Function 3 Function 1 ...

Module 4. Peptide Annotator

1. Peptide Direct to OTF from MAG Workflow

These peptide results use metagenome-assembled genomes (MAGs) as the reference database for protein searches, such as DIA-NN, MetaLab-MAG, MetaLab-DIA, and other workflows that use MAG databases like MGnify or custom MAG databases.

peptide2taxafunc

Required inputs:

Genome Selection Modes

Peptide Direct to OTF has three genome-selection modes:

MetaUmbra scoring currently requires a tab-separated peptide table. When a DIA-NN parquet file is selected, MetaX first prepares a temporary tab-separated peptide table in metax_temp before running MetaUmbra.

DIA-NN Parquet Preparation

When the input is a DIA-NN parquet file, MetaX reads only the required columns and pivots the long-format table into a direct-to-OTF peptide table:

2. MetaUmbra Unit-Specific Direct-to-OTF Annotation

MetaX can consume a MetaUmbra unit_specific_manifest.json as the preferred backend interface for unit-specific OTF annotation. In this mode, MetaX uses sample_columns from each analysis unit to split the peptide intensity table, and uses genome_ids_q005 or genome_ids_q001 to restrict peptide-to-protein mapping per unit. If --genome-threshold is not provided, the manifest default_genome_threshold is used.

This backend is additive to the normal/global Peptide Direct to OTF workflow. When unit-specific mode is disabled, MetaX uses the selected normal mode: MetaUmbra genome scoring, a user-provided genome list, or MetaUmbra scoring-only output. Unit-specific mode does not run the normal global genome-selection path; each analysis unit receives its own genome list directly from the MetaUmbra manifest.

The unit-specific distinct-genome filter defaults to 0, so MetaX trusts the manifest-selected genome list. Set --distinct-genome-threshold to a value greater than 0 only when you want an additional MetaX-side filter requiring that many distinct peptides per genome after mapping.

Sample columns are matched from manifest sample_columns to peptide-table columns in this order: exact name, Intensity_ prefix, configured output prefix, configured input prefix, stripped Intensity_, stripped output prefix, stripped input prefix, leading underscores removed, and raw-file basename without .raw, .mzML, or .mzXML. Use --input-sample-col-prefix for inputs such as LFQ intensity sample_1.

The merged unit-specific OTF table includes analysis_unit_id and the original Sequence column. MetaX internally derives the unit-specific peptide evidence ID as analysis_unit_id + "||" + Sequence when downstream analysis needs a unique peptide identity; UnitSpecificSequence is not written by default. Do not deduplicate unit-specific output by Sequence alone. Downstream final OTF identity remains Taxon + Function.

In the GUI, select the MetaUmbra unit_specific_manifest.json and genome threshold in the main Peptide Direct to OTF window. The Unit-specific Settings dialog does not select a separate manifest or threshold; it configures sample-column matching behavior and missing/empty unit handling, and validates the selected manifest against the current peptide table when possible. Unit-specific mode disables the legacy global genome scoring controls, and the duplicate peptide handling selector still applies. A manual manifest builder is not implemented yet.

Unit-specific annotation accepts either a wide peptide-intensity table with one sample intensity column per manifest sample or a long-format DIA-NN parquet containing Run, Stripped.Sequence, and either Precursor.Normalised or Precursor.Quantity. Long-format parquet input is pivoted automatically, and common raw-file suffixes such as .raw, .mzML, and .mzXML are ignored when matching Run values to manifest samples.

The default unit-specific execution path is disk-backed. Per-unit temporary files are written under <output_stem>_artifacts/per_unit/unit_otf/, merged into the final OTF table by streaming append, and then cleaned up. The final artifacts include:

For downstream analysis, unit-specific public count columns use these meanings:

Example:

metax-annotate \
  --unit-specific \
  --peptide-table report.tsv \
  --unit-specific-manifest unit_specific_manifest.json \
  --genome-threshold q0.05 \
  --taxafunc-db MetaX_taxafunc.db \
  --digested-genome-folders digested_genomes/ \
  --output OTF_unit_specific.tsv \
  --peptide-col Sequence \
  --input-sample-col-prefix "LFQ intensity " \
  --duplicate-peptide-handling-mode sum \
  --n-jobs 4

3. Results from MaxQuant Workflow

These peptide results come from the MetaLab 2.3 MaxQuant workflow.

peptide2taxafunc_tab2_1 peptide2taxafunc_tab2_2


Developer Tools

Auto OTF report

The auto report writes a self-contained MetaX_Report folder when an output parent is selected in the GUI. Existing non-empty report directories are rejected unless Overwrite is enabled, which prevents outputs from different runs being mixed.

Group-vs-control testing uses limma via InMoose by default on log2(x + 1)-transformed abundance. Zero abundance remains numeric zero during limma preprocessing. The legacy GUI Dunnett workflow remains available by setting statistics.diff_method: dunnett or using --diff-method dunnett.

The effective configuration is saved as config_used.yaml. Static figure output defaults to 300 DPI PNG and can include editable-text SVG/PDF:

statistics:
  diff_method: limma
report:
  figure_formats: [png, svg, pdf]
  dpi: 300

Equivalent CLI options are --diff-method, --figure-formats, and --dpi. The report home page identifies the main taxa level and function column, lists other combinations as extended results, and shows optional analysis-unit metadata when analysis_unit_id or a compatible unit column is present.

show_console

Enjoy MetaX

If you have any issues or suggestions, please open a new issue on GitHub.