GNM (pronounced genome) is a state-of-the-art parametric 3D statistical model of the human head developed by Google. Learned from a large dataset of high-resolution 3D scans, GNM provides fine-grained, disentangled control over facial identity, expressions, and head pose, complete with controllable internal anatomy (eyeballs, teeth, and tongue). The model is released under the Apache 2.0 permissive license, suitable for both academic research and commercial applications.
Model Overview GNM (pronounced genome /ˈdʒiː.noʊm/, in reference to the human genome) is a state-of-the-art parametric 3D statistical model of the human head developed by Google. Learned from a large dataset of high-resolution 3D scans, GNM provides fine-grained, disentangled control over facial identity, expressions, and head pose, complete with controllable internal anatomy (eyeballs, teeth, and tongue). 3D Morphable Models (3DMMs) are widely used across computer vision, computer graphics, and generative AI for representing human geometry and appearance. GNM introduces a state-of-the-art parametric representation of the human head accompanied by multi-framework backend support and semantic parameter sampling. The model is released under the Apache 2.0 permissive license, suitable for both academic research and commercial applications. Key Features & Anatomy Detailed 3D Face Geometry: Generates a dense 3D mesh consisting of the cranial/facial skin, eyes, teeth & gums, and tongue. Controllable Anatomy: Skin: Full cranial, facial, and neck coverage. Eyes: Controllable internal eyeball structures (sclera, pupil, iris) and external cornea. Teeth & Gums: Articulated upper and lower dentition. Tongue: Articulated tongue surface supporting inner-mouth deformation and speech articulations. Disentangled Parameter Spaces: Full orthogonal control over: Identity: Controls subject-specific facial features and anatomical proportions (head, eyeball, teeth). Expression: Animates the face with a rich set of blendshapes across the eyes, lower face, tongue, and iris. Head Pose & Joint Rotations: Controls kinematic rotation of the neck and bilateral eyeball gaze (axis-angle format). Translation: Controls global 3D Cartesian positioning. Multi-Topology UV Mapping: Structured UV layouts provided for both quad (quad_uvs, shape [Q, 4, 2]) and triangulated (triangle_uvs, shape [T, 3, 2]) topologies across five logical regions (skin, upper teeth/gums, lower teeth/gums, tongue, eye interior, and eye exterior). Multi-Framework Backend Support: Native support for NumPy, JAX, PyTorch, and TensorFlow via a unified GNM interface. Semantic Parameter Sampling: Compatible with pre-trained IdentitySampler and ExpressionSampler models for generating identity and expression parameters from high-level labels. Model Organization Structure Both Hugging Face Hub and Kaggle Models follow an identical organization architecture: 1 Model / Repository per MAJOR Version: Hugging Face Hub: google/gnm-v3 Kaggle Models: google/gnm-v3 Variants for MINOR Versions: Hugging Face Hub: Each minor version is organized into version subdirectories (e.g. v3_0/gnm_head.npz), with the root README.md documenting the major release and version-specific README.md files documenting each version directory. Kaggle Models: Each minor version is published as a distinct model variation under the other framework (e.g. google/gnm-v3/other/gnm_head_v3_0).
Apache 2.0
vision foundation model, feature backbone
PyTorch
Open
Science, Technology and Research
25/09/26 09:32:21
0
Apache 2.0
© 2026 - Copyright AIKosh. All rights reserved.