Q (haiku):
I'm importing a study into my own cBioPortal instance and the validator says some genes are not known to cBioPortal. What does that mean?
The validator checks every Hugo_Symbol / Entrez_Gene_Id in your data files against the gene table loaded in your instance (built from HGNC/NCBI). "Gene not known" means the symbol or ID isn't in that table — usually an outdated symbol, an alias, or a typo.
Continue this conversation: reply to the user's new message.
what if the genes, like AC002511.1 or LOC100506271, just aren't in HGNC or NCBI at all?
═ answer (success)
Those are non-coding RNA genes (lncRNAs, pseudogenes) or locusIDs that HGNC and NCBI either don't curate, or have deprecated. cBioPortal's gene table is built from HGNC/NCBI and doesn't include every sequence identifier.
You have a few options:
1. **Map to approved symbols** — Check if NCBI or Ensembl have a current approved symbol or Entrez ID for that locus. For example, `LOC100506271` may have an older Entrez ID that maps to a current gene name.
2. **Use Ensembl gene IDs** — If your data has Ensembl identifiers (ENSG*), you can sometimes cross-reference those to find an approved symbol. Tools like `biomaRt` (R) or `Ensembl REST API` can help.
3. **Filter them out** — If these genes are sparse or not central to your study, remove them from the import file. The validator will pass, and you won't lose the rest of your data.
4. **Add them to the gene table manually** (advanced) — If you have admin access to your instance, you can edit the gene table (`GENE` in your database) to add custom entries. This is rarely needed and requires careful setup.
5. **Check the validator logs** — The import tool usually lists which specific genes failed. That can help you decide if they're critical to keep.
What's your import file format (MAF, segment, expression)? And are these genes in many samples, or just a few outliers?