Q (haiku):
I'm importing a study into my own cBioPortal instance and the validator says some genes are not known to cBioPortal. What does that mean?
The validator checks every Hugo_Symbol / Entrez_Gene_Id in your data files against the gene table loaded in your instance (built from HGNC/NCBI). "Gene not known" means the symbol or ID isn't in that table — usually an outdated symbol, an alias, or a typo.
Continue this conversation: reply to the user's new message.
what if the genes, like AC002511.1 or LOC100506271, just aren't in HGNC or NCBI at all?
═ answer (success)
Those are non-coding RNA genes or pseudogenes that HGNC/NCBI either don't cover comprehensively or classify differently. Here's what's happening and how to handle it:
## Why They're Not in the Standard Gene Tables
- **AC002511.1, LOC100506271** — these are NCBI RefSeq identifiers for long non-coding RNAs (lncRNAs) or predicted genes that don't have official HUGO symbols
- HGNC focuses on protein-coding genes and well-characterized non-coding RNAs; many predicted or poorly characterized genes fall outside that scope
- NCBI includes them but may not assign stable HUGO symbols
## Solutions for Your Import
**Option 1: Map to HUGO symbols (recommended)**
- Query Ensembl or NCBI directly to see if these have HGNC-approved symbols
- Example: AC002511.1 might map to an HGNC lncRNA name like `LINC00000`
- Update your file with the official symbol if one exists
**Option 2: Use Entrez Gene IDs instead**
- If you have the Entrez ID (e.g., 100506271), the validator may accept it as the primary identifier
- cBioPortal can resolve Entrez IDs even when the symbol is ambiguous
**Option 3: Add them to your gene table**
- Edit the `cancer_gene_list.txt` file in your instance (or equivalent gene reference) to include these entries with their Entrez IDs and a descriptive name
- The validator will then recognize them
**Option 4: Filter them out**
- If these genes aren't clinically important for your study, exclude them from your import file and keep only genes cBioPortal recognizes
## Which Approach?
- **If they're actually mutated/altered in your data:** use Option 1 or 3 — map them properly so they're meaningful in analysis
- **If they're rare or background noise:** Option 4 is fine
- **For lncRNAs with clinical relevance:** Option 3 (manually add to your gene list) is most reliable
What's the source of your data? (VEP, MAF, a custom pipeline?) That context might clarify whether these genes should be kept or are annotation artifacts.