NG6 to NGL-Bi

After more than 15 years of NG6 use, we moved to NGL-Bi for data and quality control access. NGL-Bi is better suited for metadata management and will allow to make data submission easier. For sequencing runs done on GeT core facility since 01/01/2025, you can also access a tsv metadata file for ENA experiment on NGL-Bi.

We also change our data storage policy : all raw data produced by GeT core facility will be stored for only 3 months, after data releasing, for all projects. Depending on your project and the production date, there are different use cases :

  • If you have a partnership project with GeT-PlaGe (data stored for 2 years after production), or you have a contract with Bioinfo GenoToul core facility to extend the storage period of your data on NG6 :
    • If you don't plan to produce new data with GeT-PlaGe in the coming months : your data can stay on NG6. They will be erased at the end of the retention period.
    • If you plan to produce new data with GeT-PlaGe in the coming months : after interaction with Bioinfo genotoul core facility, a project space will be created for you where we will move your data from NG6 (our own reponsability) and where you will be able to store new data copied from NGL-Bi, as of 01/01/2025. Contact for any question : ng6-support(at)groupes.renater.fr
  • For all other projects
    • Your data will be deleted 3 months after their release, from NG6 or from NGL-Bi

NG6 will be maintained until 12/31/2026.

I forgot my username and password. How can I recover them?

You can access your username and change your password using this web interface.

Who can access data of my project?

Each user you defined at project creation are linked to your project. These persons have personal credentials to access NGL-Bi and can access data.

An email is automatically sent to all users linked to a project as soon as data is delivered after internal validation.

You can forward this email to other people so they can access your project data. Be aware that, doing this, gives access to all data produced for your project. It is not possible to share only a specific analysis.

In all cases, we recommend you to check the data integrity right after download, using md5sum.

What are the adapter sequences used with Illumina library (including AVITI-converted)?

When the library insert length is smaller than the sequencing read length, you can find the sequencing adapaters the end of your reads.

You can find the different sequencing adapters below :

  • TruSeq Adapter - Read 1 : "GATCGGAAGAG-CACACGTCTGAACTCCAGTC"
  • Illumina single End PCR Primer - Read 2 : "GATCGGAAGAG-CGTCGTGTAGGGAAAGAGTG"
  • Nextera  transposase sequence - Read 1 : 5’ TCGTCGGCAGCGTC-AGATGTGTATAAGAGACAG
  • Nextera  transposase sequence - Read 2 : 5’ GTCTCGTGGGCTCGG-AGATGTGTATAAGAGACAG

You can use cutadapt to remove the adapter sequences, specifying the following sequences : 

  • Nextera index : CTGTCTCTTATACACATCT
  • Illlumina Index : AGATCGGAAGAGC
What are the adapter sequences used with Oxford Nanopore technology?

Nanopore reads can contain adapters and barcodes, for multiplexed runs.

Dorado can help you remove these ones.

Check the sequencing kit name used for library preparation in ToulligQC or sequencing report at analysis level in NGL-Bi.

  • SQK-LSK114 adapter sequences
    • Top strand: 5'-TTTTTTTTCCTGTACTTCGTTCAGTTACGTATTGCT-3’
    • Bottom strand: 5’-GCAATACGTAACTGAACGAAGTACAGG-3’
  • SQK-NBD114-24 adapter sequences
    • Top strand: 5'-TTTTTTTTCCTGTACTTCGTTCAGTTACGTATTGCT-3'
    • Bottom strand: 5'-ACGTAACTGAACGAAGTACAGG-3'
  • Barcode flanking sequences
    • Forward sequence: 5' - AAGGTTAA - barcode - CAGCACCT - 3'
    • Reverse sequence: 5' - GGTGCTG - barcode - TTAACCTTAGCAAT - 3'

You can find official information on https://nanoporetech.com/document/chemistry-technical-document#barcode-sequences.

Native Barcoding Expansion 1-12

Component Sequence
NB01 CACAAAGACACCGACAACTTTCTT
NB02 ACAGACGACTACAAACGGAATCGA
NB03 CCTGGTAACTGGGACACAAGACTC
NB04 TAGGGAAACACGATAGAATCCGAA
NB05 AAGGTTACACAAACCCTGGACAAG
NB06 GACTACTTTCTGCCTTTGCGAGAA
NB07 AAGGATTCATTCCCACGGTAACAC
NB08 ACGTAACTTGGTTTGTTCCCTGAA
NB09 AACCAAGACTCGCTGTGCCTAGTT
NB10 GAGAGGACAAAGGTTTCAACGCTT
NB11 TCCATTCCCTCCGATAGATGAAAC
NB12 TCCGATTCTGCTTCTTTCTACCTG

Native Barcoding Expansion 13-24

Component Sequence
NB13 TCACACGAGTATGGAAGTCGTTCT
NB14 TCTATGGGTCCCAAGAGACTCGTT
NB15 CAGTGGTGTTAGCGAGGTAGACCT
NB16 AGTACGAACCACTGTCAGTTGACG
NB17 ATCAGAGGTACTTTCCTGGAGGGT
NB18 GCCTATCTAGGTTGTTGGGTTTGG
NB19 ATCTCTTGACACTGCACGAGGAAC
NB20 ATGAGTTCTCGTAACAGGACGCAA
NB21 TAGAGAACGGACAATGAGAGGCTC
NB22 CGTACTTTGATACATGGCAGTGGT
NB23 CGAGGAGGTTCACTGGGTAGTAAG
NB24 CTAACCCATCATGCAGAACTATGC

Barcode flanking sequences

In addition to the barcodes themselves, the barcoding primers have additional flanking sequences, which add an extra level of context during analysis. The flanking sequences differ depending on the kit:

Native Barcoding Expansion 1-12 (EXP-NBD104) and Native Barcoding Expansion 13-24

5' - AAGGTT - barcode - ACAGCACCT - 3'

How long are my data stored?

The data retention period is set to 3 months, after its validation.

You will receive a reminder email a month before your data exceed the retention limit.

Why do I have several fastq files per sample ? And how to deal with these?

Depending on the amount of data required for your project, it may be necessary to sequence your samples multiple times. We provide the data sequencing lane by sequencing lane, so you may receive multiple fastq files per sample. To run your analyses on your entire data set, and depending on your analyses, you could want to merge the different fastq files per sample name. After retrieving your data, fastq files should be grouped into directories with the following tree structure:

					PROJECT/SEQUENCING_LANE_DATA_1/  
					   ├── date_sample-1_L001_R1.fastq.gz  
					   ├── date_sample-1_L001_R2.fastq.gz  
					   ├── date_sample-2_L001_R1.fastq.gz  
					   └── date_sample-2_L001_R2.fastq.gz  

					PROJECT/SEQUENCING_LANE_DATA_2/  
					   ├── date_sample-1_L002_R1.fastq.gz  
					   ├── date_sample-1_L002_R2.fastq.gz  
					   ├── date_sample-2_L002_R1.fastq.gz  
					   └── date_sample-2_L002_R2.fastq.gz  
					

Snippet you can use to merge files (on Linux/Unix terminal): Please review the data structure requirements before use it (see above - before october 2025, the fastq files names were sample-1_L002_R1.fastq.gz, without the date on first part). This code snippet is provided for informational purposes only and may require modifications before being applied to your data.

					PROJECT=__input_path__
					MERGE_DIR=__output_path__
					mkdir -p $MERGE_DIR
					for sample in `find $PROJECT -type f -name '*.fastq.gz'`; do
					    sampleBasename=$(basename $sample)
					    sampleName=$(echo $sampleBasename | cut -d_ -f2)
					    readNumberWithExt=$(echo $sampleBasename | cut -d_ -f5)
					    cat $sample >> ${MERGE_DIR}/${sampleName}_${readNumberWithExt}
					done
					

I received Undetermined fastq files from an AVITI sequencing. How to perform demultiplexing from i7 and/or i5 indexes ?

Undetermined fastq files contain only reads with no match regarding the provided index sequences at demultiplex step. You could be interested in analyzing these files in some circumstances.

To perform demultiplexing on Undetermined fastq files, you can use the fastq_demux tool (https://github.com/Molmed/fastq_demux).

This tool only uses i7 and/or i5 indexes, not UMI inside read sequences.

To use this tool, you must provide a sampleSheet which is a tab-separated file containing a sample name per row associated with its index sequence(s).

To get every index sequences from an undertermined fastq file, run something like:

zcat my_undetermined_fastq_file.fastq.gz | awk 'NR % 4 == 1 {split($0,a,":"); print a[length(a)]}' | sort | uniq -c | sort -nr > index_count.tsv

NB : To execute this command, you can use any Undetermined file (R1, R2, I1 or I2); the result will be exactly the same. However, using one of the two index files will make the command faster.

To maximize the number of output sequences, we recommend using 1 mismatch (--mismatches 1) with 6-bases indexes and 2 mismatches (--mismatches 2) with longer indexes.

For more details, have a look at the tool documentation on GitHub. Please note that this is not a tool developed within GeT plateform, but only a recommendation. We can help you use the tool if needed, but we will not provide support in case of dysfunction. If necessary, contact the developer directly opening an issue on GitHub.

What do the files I received correspond to for AVITI (Element Biosciences platform)?

FASTQ

The FASTQ files provided by the GeT platforms contain demultiplexed sequences and quality scores. These files are compressed in gzip format.

Files with the tags R1 or R2 are read sequence files.

Files with the tags I1 or I2 are index sequence files.

Undetermined files contain undemultiplexed data. This type of file is provided only in exceptional circumstances.

What do the files I received correspond to for PromethION or GridION (Oxford Nanopore Technologies platforms)?

FASTQ

The FASTQ files provided by the GeT platforms contain demultiplexed sequences and quality scores. These files are compressed in gzip format. Unclassified files contain undemultiplexed reads.

Sequencing Summary

There are large tabular files containing extensive sequencing metadata and metrics, including Adaptive Sampling decisions, where applicable. They are compressed in gzip format.

ModBAM

This type of file contains epigenetic marks (such as methylation).

BED

This is a tabular text file summarizing the contents of the ModBAM file.

POD5

Raw Data Files: These files contain the metadata and each electrical potential difference acquired during sequencing. This large file is provided only upon request. It is only useful if you wish to basecall again the sequencing data or extract epigenetic marks by your own. Processing it requires significant dedicated GPU resources.

AS_decisions.csv

Only release when Adaptive Sampling process is enable during sequencing. For each read, this file indicates whether it was accepted or rejected, and whether the sequencer was able to implement that decision.

What do the files I received correspond to for VEGA (Pacific BioSciences platform)?

General principle: “BAM everywhere”

With VEGA, BAM is the central and universal data format. BAM files are used as:

  • the primary output of the sequencing run,
  • the input to downstream analysis tools (mapping, assembly, Iso-Seq, modification detection),

BAM format: general definition

BAM (Binary Alignment Map) is the compressed binary representation of the SAM (Sequence Alignment/Map) format.

Despite its historical name, a BAM file does not necessarily contain aligned reads.

A BAM file stores:

  • nucleotide sequences,
  • per-base quality values,
  • optional metadata encoded as tags,
  • global metadata stored in the file header.

Whether reads are aligned or unaligned depends on the producing application, not on the format itself.

VEGA specificity: HiFi reads as primary data

On VEGA, the first-class data type is the HiFi read :

  • HiFi reads are consensus reads (CCS) generated from multiple passes of the same DNA molecule,
  • Only reads with a predicted accuracy QV ≥ 20 are emitted,
  • HiFi reads are stored in unaligned BAM files by default.

Importantly, raw subreads are no longer accessible:

  • Since SMRT Link 11.0 on Sequel IIe, subreads are no longer exposed,
  • Subreads are not available at all on Revio and VEGA.

What is means : downstream analyses rely entirely on the HiFi consensus reads and on the metadata preserved in BAM tags. Recomputing the consensus from subreads is not possible.

BAM files in analysis workflows

  • Alignment (mapping) tools take unaligned BAM files as input and produce aligned BAM files,
  • During alignment, all existing tags and headers are preserved,
  • Secondary analysis applications (e.g. Iso-Seq, assembly, polishing, modification detection) typically consume aligned or unaligned BAM files and generate new BAM files as output.

Unaligned PacBio HiFi reads are grouped by ZMW hole number, but are not sorted by coordinate or by hole number.

BAM header content (PacBio extensions)

In addition to standard SAM/BAM header fields, PacBio encodes run- and file-level metadata in the header, following documented conventions.

Key header records include :

  • HD: format version and sorting information.
  • RG (Read Group): identifies logical groups of reads and includes:
    • run and platform information (ID, PL, PM, PU),
    • sample and library information (SM, LB),
    • barcode sequence (BC),
    • descriptive metadata (DS).
  • PG (Program Group): records software tools and versions that produced or modified the BAM file.

These headers allow traceability of data provenance across analysis steps.

Read records and naming conventions

Each BAM record corresponds to one read.

PacBio uses a standardized read naming convention (QNAME):

{movieName}/{holeNumber}/ccs

In some cases, segmented CCS reads may be present. These represent partial CCS segments and are identified as:

{movieName}/{holeNumber}/ccs
{movieName}/{holeNumber}/ccs/fwd
{movieName}/{holeNumber}/ccs/rev

Segmented reads can affect downstream processing and should be handled explicitly by analysis tools.

HiFi per-read tags

PacBio HiFi BAM files include per-read metadata encoded as SAM tags, notably:

  • np: number of polymerase passes used to generate the CCS read.
  • rq: predicted read quality (HiFi accuracy estimate).

Kinetic information

  • By default, average kinetic information may be available at the read level for specific applications (Microbial assembly).
  • Per-base kinetic information can be produced, but this significantly increases data volume (approximately 5× the storage of a standard run).

Availability of kinetic data depend on the selected application and instrument configuration.

Base modification information

PacBio encodes base modification data using standard SAM tags:

  • MM: base modification types and positions (e.g. methylation contexts),
  • ML: modification likelihoods or probabilities.

These tags enable downstream detection and analysis of DNA base modifications.

Iso-Seq–specific tags

For Iso-Seq applications, additional CCS read tags are included to support transcript annotation and downstream Iso-Seq analyses, following PacBio specifications.