After more than 15 years of NG6 use, we moved to NGL-Bi for data and quality control access. NGL-Bi is better suited for metadata management and will allow to make data submission easier. For sequencing runs done on GeT core facility since 01/01/2025, you can also access a tsv metadata file for ENA experiment on NGL-Bi.
We also change our data storage policy : all raw data produced by GeT core facility will be stored for only 3 months, after data releasing, for all projects. Depending on your project and the production date, there are different use cases :
NG6 will be maintained until 12/31/2026.
You can access your username and change your password using this web interface.
Each user you defined at project creation are linked to your project. These persons have personal credentials to access NGL-Bi and can access data.
An email is automatically sent to all users linked to a project as soon as data is delivered after internal validation.
You can forward this email to other people so they can access your project data. Be aware that, doing this, gives access to all data produced for your project. It is not possible to share only a specific analysis.
In all cases, we recommend you to check the data integrity right after download, using md5sum.
When the library insert length is smaller than the sequencing read length, you can find the sequencing adapaters the end of your reads.
You can find the different sequencing adapters below :
You can use cutadapt to remove the adapter sequences, specifying the following sequences :
Nanopore reads can contain adapters and barcodes, for multiplexed runs.
Dorado can help you remove these ones.
Check the sequencing kit name used for library preparation in ToulligQC or sequencing report at analysis level in NGL-Bi.
You can find official information on https://nanoporetech.com/document/chemistry-technical-document#barcode-sequences.
Native Barcoding Expansion 1-12
| Component | Sequence |
|---|---|
| NB01 | CACAAAGACACCGACAACTTTCTT |
| NB02 | ACAGACGACTACAAACGGAATCGA |
| NB03 | CCTGGTAACTGGGACACAAGACTC |
| NB04 | TAGGGAAACACGATAGAATCCGAA |
| NB05 | AAGGTTACACAAACCCTGGACAAG |
| NB06 | GACTACTTTCTGCCTTTGCGAGAA |
| NB07 | AAGGATTCATTCCCACGGTAACAC |
| NB08 | ACGTAACTTGGTTTGTTCCCTGAA |
| NB09 | AACCAAGACTCGCTGTGCCTAGTT |
| NB10 | GAGAGGACAAAGGTTTCAACGCTT |
| NB11 | TCCATTCCCTCCGATAGATGAAAC |
| NB12 | TCCGATTCTGCTTCTTTCTACCTG |
Native Barcoding Expansion 13-24
| Component | Sequence |
|---|---|
| NB13 | TCACACGAGTATGGAAGTCGTTCT |
| NB14 | TCTATGGGTCCCAAGAGACTCGTT |
| NB15 | CAGTGGTGTTAGCGAGGTAGACCT |
| NB16 | AGTACGAACCACTGTCAGTTGACG |
| NB17 | ATCAGAGGTACTTTCCTGGAGGGT |
| NB18 | GCCTATCTAGGTTGTTGGGTTTGG |
| NB19 | ATCTCTTGACACTGCACGAGGAAC |
| NB20 | ATGAGTTCTCGTAACAGGACGCAA |
| NB21 | TAGAGAACGGACAATGAGAGGCTC |
| NB22 | CGTACTTTGATACATGGCAGTGGT |
| NB23 | CGAGGAGGTTCACTGGGTAGTAAG |
| NB24 | CTAACCCATCATGCAGAACTATGC |
Barcode flanking sequences
In addition to the barcodes themselves, the barcoding primers have additional flanking sequences, which add an extra level of context during analysis. The flanking sequences differ depending on the kit:
5' - AAGGTT - barcode - ACAGCACCT - 3'
The data retention period is set to 3 months, after its validation.
You will receive a reminder email a month before your data exceed the retention limit.
Depending on the amount of data required for your project, it may be necessary to sequence your samples multiple times. We provide the data sequencing lane by sequencing lane, so you may receive multiple fastq files per sample. To run your analyses on your entire data set, and depending on your analyses, you could want to merge the different fastq files per sample name. After retrieving your data, fastq files should be grouped into directories with the following tree structure:
PROJECT/SEQUENCING_LANE_DATA_1/ ├── date_sample-1_L001_R1.fastq.gz ├── date_sample-1_L001_R2.fastq.gz ├── date_sample-2_L001_R1.fastq.gz └── date_sample-2_L001_R2.fastq.gz PROJECT/SEQUENCING_LANE_DATA_2/ ├── date_sample-1_L002_R1.fastq.gz ├── date_sample-1_L002_R2.fastq.gz ├── date_sample-2_L002_R1.fastq.gz └── date_sample-2_L002_R2.fastq.gz
Snippet you can use to merge files (on Linux/Unix terminal): Please review the data structure requirements before use it (see above - before october 2025, the fastq files names were sample-1_L002_R1.fastq.gz, without the date on first part). This code snippet is provided for informational purposes only and may require modifications before being applied to your data.
PROJECT=__input_path__
MERGE_DIR=__output_path__
mkdir -p $MERGE_DIR
for sample in `find $PROJECT -type f -name '*.fastq.gz'`; do
sampleBasename=$(basename $sample)
sampleName=$(echo $sampleBasename | cut -d_ -f2)
readNumberWithExt=$(echo $sampleBasename | cut -d_ -f5)
cat $sample >> ${MERGE_DIR}/${sampleName}_${readNumberWithExt}
done
Undetermined fastq files contain only reads with no match regarding the provided index sequences at demultiplex step. You could be interested in analyzing these files in some circumstances.
To perform demultiplexing on Undetermined fastq files, you can use the fastq_demux tool (https://github.com/Molmed/fastq_demux).
This tool only uses i7 and/or i5 indexes, not UMI inside read sequences.
To use this tool, you must provide a sampleSheet which is a tab-separated file containing a sample name per row associated with its index sequence(s).
To get every index sequences from an undertermined fastq file, run something like:
zcat my_undetermined_fastq_file.fastq.gz | awk 'NR % 4 == 1 {split($0,a,":"); print a[length(a)]}' | sort | uniq -c | sort -nr > index_count.tsv
NB : To execute this command, you can use any Undetermined file (R1, R2, I1 or I2); the result will be exactly the same. However, using one of the two index files will make the command faster.
To maximize the number of output sequences, we recommend using 1 mismatch (--mismatches 1) with 6-bases indexes and 2 mismatches (--mismatches 2) with longer indexes.
For more details, have a look at the tool documentation on GitHub. Please note that this is not a tool developed within GeT plateform, but only a recommendation. We can help you use the tool if needed, but we will not provide support in case of dysfunction. If necessary, contact the developer directly opening an issue on GitHub.
FASTQ
The FASTQ files provided by the GeT platforms contain demultiplexed sequences and quality scores. These files are compressed in gzip format.
Files with the tags R1 or R2 are read sequence files.
Files with the tags I1 or I2 are index sequence files.
Undetermined files contain undemultiplexed data. This type of file is provided only in exceptional circumstances.
FASTQ
The FASTQ files provided by the GeT platforms contain demultiplexed sequences and quality scores. These files are compressed in gzip format. Unclassified files contain undemultiplexed reads.
Sequencing Summary
There are large tabular files containing extensive sequencing metadata and metrics, including Adaptive Sampling decisions, where applicable. They are compressed in gzip format.
ModBAM
This type of file contains epigenetic marks (such as methylation).
BED
This is a tabular text file summarizing the contents of the ModBAM file.
POD5
Raw Data Files: These files contain the metadata and each electrical potential difference acquired during sequencing. This large file is provided only upon request. It is only useful if you wish to basecall again the sequencing data or extract epigenetic marks by your own. Processing it requires significant dedicated GPU resources.
AS_decisions.csv
Only release when Adaptive Sampling process is enable during sequencing. For each read, this file indicates whether it was accepted or rejected, and whether the sequencer was able to implement that decision.
General principle: “BAM everywhere”
With VEGA, BAM is the central and universal data format. BAM files are used as:
BAM format: general definition
BAM (Binary Alignment Map) is the compressed binary representation of the SAM (Sequence Alignment/Map) format.
Despite its historical name, a BAM file does not necessarily contain aligned reads.
A BAM file stores:
Whether reads are aligned or unaligned depends on the producing application, not on the format itself.
VEGA specificity: HiFi reads as primary data
On VEGA, the first-class data type is the HiFi read :
Importantly, raw subreads are no longer accessible:
What is means : downstream analyses rely entirely on the HiFi consensus reads and on the metadata preserved in BAM tags. Recomputing the consensus from subreads is not possible.
BAM files in analysis workflows
Unaligned PacBio HiFi reads are grouped by ZMW hole number, but are not sorted by coordinate or by hole number.
BAM header content (PacBio extensions)
In addition to standard SAM/BAM header fields, PacBio encodes run- and file-level metadata in the header, following documented conventions.
Key header records include :
These headers allow traceability of data provenance across analysis steps.
Read records and naming conventions
Each BAM record corresponds to one read.
PacBio uses a standardized read naming convention (QNAME):
{movieName}/{holeNumber}/ccs
In some cases, segmented CCS reads may be present. These represent partial CCS segments and are identified as:
{movieName}/{holeNumber}/ccs
{movieName}/{holeNumber}/ccs/fwd
{movieName}/{holeNumber}/ccs/rev
Segmented reads can affect downstream processing and should be handled explicitly by analysis tools.
HiFi per-read tags
PacBio HiFi BAM files include per-read metadata encoded as SAM tags, notably:
Kinetic information
Availability of kinetic data depend on the selected application and instrument configuration.
Base modification information
PacBio encodes base modification data using standard SAM tags:
These tags enable downstream detection and analysis of DNA base modifications.
Iso-Seq–specific tags
For Iso-Seq applications, additional CCS read tags are included to support transcript annotation and downstream Iso-Seq analyses, following PacBio specifications.