Q (sonnet): I'm importing a study into my own cBioPortal instance and the validator says some genes are not known to cBioPortal. What does that mean? The validator checks every Hugo_Symbol / Entrez_Gene_Id in your data files against the gene table loaded in your instance (built from HGNC/NCBI). "Gene not known" means the symbol or ID isn't in that table — usually an outdated symbol, an alias, or a typo. Continue this conversation: reply to the user's new message. what if the genes, like AC002511.1 or LOC100506271, just aren't in HGNC or NCBI at all? ═ answer (success) Good question — those are real, valid gene identifiers, just not conventional protein-coding gene symbols with stable HGNC entries. - **AC002511.1** — this is a clone-based Ensembl/GenBank identifier (accession.version format), typically used for a novel transcript or lncRNA locus before (or instead of) it gets an official HGNC symbol. It reflects the genomic clone the sequence was originally annotated on. - **LOC100506271** — this is an NCBI-generated placeholder symbol, automatically assigned by NCBI's annotation pipeline for a locus that has an Entrez Gene ID but no approved HGNC gene symbol yet. Both are legitimate, stable identifiers in their respective source databases (Ensembl/GenBank and NCBI, respectively), but they may not be present in the specific gene table snapshot your cBioPortal instance loaded, especially if that snapshot only includes HGNC-approved symbols and their aliases rather than the full Ensembl/NCBI gene sets. A few practical points, purely about how cBioPortal's import process works (not clinical or biological guidance): 1. **The validator flags, but often doesn't block import** — depending on the file type, unknown genes may just be skipped/dropped from that record rather than failing the whole file. Check the validator output carefully for whether it's a warning or an error. 2. **Your instance's gene table can be updated.** cBioPortal periodically refreshes its reference gene table (via `importGenes` scripts) from combined HGNC + NCBI Entrez Gene data dumps, which do include many `LOC*` placeholder symbols and some Ensembl clone-based IDs. If your instance was seeded from an older or trimmed gene table, re-importing a current seed file may resolve it. 3. **Entrez_Gene_Id is more robust than Hugo_Symbol** for loci like `LOC100506271`, since the numeric Entrez ID is what's actually indexed — if your file only has the symbol and no matching Entrez ID, that's a common cause of failed resolution. 4. **Clone-based accessions like AC002511.1** are less likely to be in any gene table at all (even NCBI's), since they're archival identifiers, not standardized gene symbols — you may need to map these to an approved symbol/Entrez ID yourself upstream, or exclude that record if it's not a locus you care about. This is a data-import/schema question rather than something answerable from the cBioPortal database itself, so I don't have query results to show here. If you want, I can point you to the specific cBioPortal documentation page on gene table maintenance and the `importGenes.pl` seed process — just let me know.