Abstract In challenging urban scenarios with low illumination and changing fields of view, the accuracy and stability significantly degrade for the existing multi-sensor fusion positioning methods. To address this degradation, a multi-modal tightly coupled positioning framework based on the error-state iterated Kalman filter is established, integrating thermal camera, LiDAR, and IMU. In addition, an adaptive enhancement method based on field of view perception is proposed within this framework. After data preprocessing, targeting the measurement characteristics of point cloud spatial distribution in open and narrow field of view scenarios, an adaptive factor to capture field of view features is designed. This factor dynamically adjusts the current-frame point cloud density, root voxel map resolution, and maximum iteration number in layer-by-layer updates, establishing a quantitative mapping from field of view characteristics to front-end and back-end system parameters. The proposed method is validated on both open-source and self-collected datasets in urban scenarios, exhibiting visibly better positioning accuracy and stability, with a 18.78% average reduction in absolute trajectory error over the second-best method. And the effectiveness of each module is verified through ablation experiments. The open-source code is available at: https://github.com/GNSSer-zzh/A–I-LITO .