arXiv:2608.19574v1 Announce Type: new Abstract: World action models jointly predict future visual observations and actions, whereas existing tactile-aware variants typically represent future touch as an image or latent stream without modeling the physical dependencies that organize tactile states h