Q (haiku): I'm importing a study into my own cBioPortal instance and the validator says some genes are not known to cBioPortal. What does that mean? The validator checks every Hugo_Symbol / Entrez_Gene_Id in your data files against the gene table loaded in your instance (built from HGNC/NCBI). "Gene not known" means the symbol or ID isn't in that table — usually an outdated symbol, an alias, or a typo. Continue this conversation: reply to the user's new message. what if the genes, like AC002511.1 or LOC100506271, just aren't in HGNC or NCBI at all? ═ answer (success) Those are non-coding RNA genes or pseudogenes that HGNC/NCBI either don't cover comprehensively or classify differently. Here's what's happening and how to handle it: ## Why They're Not in the Standard Gene Tables - **AC002511.1, LOC100506271** — these are NCBI RefSeq identifiers for long non-coding RNAs (lncRNAs) or predicted genes that don't have official HUGO symbols - HGNC focuses on protein-coding genes and well-characterized non-coding RNAs; many predicted or poorly characterized genes fall outside that scope - NCBI includes them but may not assign stable HUGO symbols ## Solutions for Your Import **Option 1: Map to HUGO symbols (recommended)** - Query Ensembl or NCBI directly to see if these have HGNC-approved symbols - Example: AC002511.1 might map to an HGNC lncRNA name like `LINC00000` - Update your file with the official symbol if one exists **Option 2: Use Entrez Gene IDs instead** - If you have the Entrez ID (e.g., 100506271), the validator may accept it as the primary identifier - cBioPortal can resolve Entrez IDs even when the symbol is ambiguous **Option 3: Add them to your gene table** - Edit the `cancer_gene_list.txt` file in your instance (or equivalent gene reference) to include these entries with their Entrez IDs and a descriptive name - The validator will then recognize them **Option 4: Filter them out** - If these genes aren't clinically important for your study, exclude them from your import file and keep only genes cBioPortal recognizes ## Which Approach? - **If they're actually mutated/altered in your data:** use Option 1 or 3 — map them properly so they're meaningful in analysis - **If they're rare or background noise:** Option 4 is fine - **For lncRNAs with clinical relevance:** Option 3 (manually add to your gene list) is most reliable What's the source of your data? (VEP, MAF, a custom pipeline?) That context might clarify whether these genes should be kept or are annotation artifacts.