Skip to main content

GEA

Last updated 2026/06/29Edit

GEA is the DDBJ Center public archive for functional genomics data, accepting microarray and sequencing experiments in MAGE-TAB format with E-GEAD-n accessions.

What is GEA

GEA (Genomic Expression Archive) is the DDBJ Center public archive for functional genomics data. It accepts experimental data derived from microarray and sequencing platforms, covering gene expression, epigenetics, and SNP array genotyping.

GEA follows the MIAME (microarray) and MINSEQE (sequencing) guidelines, and metadata are described in the MAGE-TAB format. GEA plays the same role at DDBJ that NCBI GEO and EBI ArrayExpress (now part of BioStudies) play at their organisations, but data are not mirrored between them and each archive uses its own accession namespace.

NOTE

If you are unsure which DDBJ service to submit your data to, the Submit Navigator walks you through the decision interactively.

Accepted data

GEA covers the following functional genomics experiments.

Experiment typeExamplesRoute
MicroarrayGene expression array, methylation array, SNP genotyping arrayMicroarray
High-throughput sequencingRNA-seq, ChIP-seq, ATAC-seq and other expression / epigenetics assaysSequencing
Single-cellscRNA-seq and other single-cell assaysSequencing
Spatial transcriptome (sequence-based)10x Genomics VisiumSequencing
Spatial transcriptome (image-based)10x Genomics Xenium, MERFISHMicroarray
Analysis using a transcriptome referenceReference-based expression quantificationSequencing

The required files differ depending on the route.

  • Microarray: Upload both raw data and processed data to GEA.
  • Sequencing: Processed (analyzed) data are required at GEA. Raw reads are submitted to DRA, and the GEA submission references the DRA submission.

WARNING

BAM / SAM / BED files alone cannot be registered as processed data. Quantitative data such as an expression matrix are required. If you only have those files, please contact GEA in advance.

IMPORTANT

A single submission cannot mix microarray and sequencing data. If you need to register both, split them into separate submissions. The maximum number of SDRF assays per submission is 1,000.

Microarray and Sequencing

GEA has two submission routes, which differ in the prerequisites and in where raw data are stored.

AspectMicroarraySequencing
Raw data locationGEADRA
Processed dataRequired at GEARequired at GEA
Prerequisite submissions to referenceBioProject + BioSampleDRA submission + BioProject (BioSample is referenced via DRA)
SDRF template sourceBioSampleDRA submission
Array DesignA-XXXX-n (existing) or upload an ADF to issue a new oneNot required
Spatial transcriptomeXenium, MERFISH, etc.Visium, etc.

Accession numbers

GEA issues three types of accessions, all assigned upon completion of curation. The accession cited in publications is the Experiment accession (E-GEAD-n).

TypePrefixExampleMeaning
ExperimentE-GEAD-E-GEAD-100The whole submission. Recorded in IDF Comment[GEAAccession]
Array designA-GEAD-A-GEAD-10New array designs issued by GEA. Existing ArrayExpress accessions such as A-AFFY-2 can also be referenced from IDF / SDRF
ProtocolP-GEAD-P-GEAD-100Each protocol defined in the IDF. Before assignment a temporary ID such as ESUB000500_Protocol_1 is used, and it is replaced once the accession is issued

A reviewer token can optionally be issued so peer reviewers can access the submission before publication.

Submission flow

  1. Sign in with your D-way account and create a new submission from the GEA submission page. An FTP upload directory is allocated for you.
  2. Reference the prerequisite submissions. For microarray, select a BioProject and a BioSample; for sequencing, select one DRA submission and a BioProject.
  3. Fill in the IDF tab with experiment-wide metadata (title / description / experiment type / design / protocol / publication / array design ref, etc.).
  4. Fill in the SDRF tab by completing the template auto-generated from the BioSample (or DRA submission) and array design, adding Material Type / Label / Factor Value / Array Data File entries with md5 values.
  5. Upload data files via FTP (raw / processed).
  6. Undergo curation and respond to any revision requests. Once complete, E-GEAD-n / A-GEAD-n / P-GEAD-n are issued.

TIP

Step-by-step wizard instructions are provided by the step cards in the Submit Navigator. This page only describes the high-level routes.

Prerequisites (MAGE-TAB)

GEA metadata are centered on the MAGE-TAB format. At minimum prepare an IDF and an SDRF, and add an ADF when registering a new array.

FileRoleNotes
IDF (Investigation Description Format)Describes the experiment as a whole (title, description, design, protocol, publication, etc.)One file
SDRF (Sample and Data Relationship Format)Maps samples to data filesFactor Value columns must be placed rightmost
ADF (Array Design Format)Defines a new array designNot needed when reusing an existing array

In addition, prepare the following beforehand.

  • D-way account: required to sign in to the GEA submission UI
  • BioProject: one BioProject is referenced by both the microarray and the sequencing route
  • BioSample: selected directly in the microarray route as the source for SDRF auto-generation (in the sequencing route it is referenced via the DRA submission)
  • DRA submission: pre-register raw reads in the sequencing route
  • Array Design: in the microarray route, reference an existing A-XXXX-n, or upload an ADF for a new design
  • Data files: raw / processed files for each assay, with md5 values

IMPORTANT

Spreadsheet-shaped files (IDF / SDRF / expression matrices, etc.) must be saved as tab-delimited .txt. .xls / .xlsx are not accepted. File names may only contain alphanumerics, _, -, and .; spaces and parentheses are not allowed.

NOTE

For two-color arrays, two samples must be linked to a single raw data file. Use the Label column in SDRF to distinguish Cy3 / Cy5.

Spatial transcriptome

The submission route for spatial transcriptome data depends on the platform.

PlatformSubmission TypeArray DesignNotes
10x Genomics VisiumSequencingNot requiredRaw read fastq/bam goes to DRA. Images (tissue_hires_image.png etc.), scale factors (scalefactors_json.json), spot positions (tissue_positions_list.csv), and the expression matrix are bundled as a tar archive and submitted as GEA processed data
10x Genomics XeniumMicroarrayA-GEAD-246Raw files such as morphology.ome.tif and processed files such as cell_feature_matrix.h5 are bundled as tar archives
MERFISH / MERSCOPEMicroarrayA-GEAD-247A dummy raw data file is submitted alongside processed data listing the identified transcripts

WARNING

Raw MERFISH images and .vzg files are not accepted by GEA. Deposit them in a generalist archive (Figshare, Zenodo, etc.). Only quantified expression data are stored in GEA.

Two-step submission with DRA

The sequencing route does not complete within GEA alone; it is a two-step submission with DRA.

flowchart LR
  A[raw read fastq/BAM]
  B[DRA submission]
  C[GEA submission]
  D[processed data]
  A -->|registered first| B
  B -->|references| C
  C --> D

When creating a GEA submission, select one DRA submission registered under your own account. The SDRF template is auto-generated from the experiment / run information of that DRA submission. Submitting sample-level processed data to GEA is strongly recommended.

IMPORTANT

GEA itself is not an INSDC member archive, but the sequencing route sits on top of the INSDC layer (DRA / BioProject / BioSample). Sequencing experiments that include raw reads always require prior submission to DRA.