During audiovisual perception, spatial information from vision and audition is combined, often producing biases such as the ventriloquist effect. While these interactions are well documented behaviourally, it remains unclear when cross-modal information begins to alter modality-specific spatial representations in the brain. Here we used cross-generalised inverted encoding modelling of electroencephalography (EEG) data to track the temporal evolution of spatial representations during a spatial ventriloquist task. Human participants (both sexes) localised audiovisual stimuli with horizontally offset auditory and visual components. Decoders trained on unisensory EEG responses were applied to audiovisual trials to estimate whether and when spatial representations of one modality were biased toward the other. Behaviourally, participants integrated cues but showed slight visual over-weighting. Neural decoding revealed robust spatial representations for both modalities, with early unisensory encoding remaining unaffected by cross-modal input. Cross-modal biases emerged only later (from ∼200 ms onwards), within a generalised representational window, and occurred sequentially: auditory representations were biased earlier than visual. These findings indicate that multisensory spatial integration arises via recurrent feedback during later processing stages, while early sensory-specific representations remain independent, providing a neural basis for the temporal dynamics underlying the ventriloquist effect. Significance statement Understanding how the brain integrates spatial information across the senses is central to theories of perception, yet the timing of cross-modal influences on modality-specific neural representations remains unclear. Using EEG and cross-generalised inverted encoding models, we tracked auditory and visual spatial representations with millisecond precision. Early spatial representations remained unisensory despite synchronous audiovisual stimulation. A shared, generalised spatial representation emerged at ∼200 ms, after which cross-modal biases appeared-first in auditory and later in visual codes. These results show that spatial integration is implemented through late, feedback-driven processes operating on initially independent sensory estimates, resolving a long-standing discrepancy between early multisensory interactions and late spatial ventriloquism.