Forest inventories are essential for monitoring forest resources at national and regional scales. Most existing approaches estimate structural attributes from airborne LiDAR point clouds and species composition from multispectral imagery using hand-crafted features. However, joint exploitation of both modalities remains underexplored, and manual feature engineering can limit performance—particularly for 3D point clouds that are heavily simplified into one-dimensional metrics. For multispectral imagery, feature engineering is increasingly replaced by geospatial foundation model (GFM) embeddings (e.g., AlphaEarth and TESSERA), which provide richer representations of Earth observation data.We propose a unified deep learning framework, based on Point Transformer v3, that jointly estimates forest structural parameters and tree species composition, by combining airborne LiDAR with georeferenced GFM embeddings. We show that learned features from the LiDAR point cloud outperform hand-crafted features, and that mid-level fusion of those LiDAR features with GFM embeddings further improves species classification performance, while structural attributes are primarily driven by LiDAR data.The results highlight the potential of GFM embeddings in combination with learned LiDAR features, for scalable and generalizable forest attribute mapping.