Music-driven 3D dance motion generation is a critical and challenging task in cross-modal content creation, aiming to synthesize physically plausible human dance sequences synchronized with input music in terms of both rhythm and style. Although diffusion-based methods have made significant progress, existing approaches typically model coupled spatiotemporal dimensions and struggle to simultaneously capture the beat structure of music and the complex dynamics of human poses with precision. To address this issue, we propose a novel multiscale spatiotemporal hybrid attention framework. The core innovation lies in decoupling the spatiotemporal complexity of dance generation: for temporal modeling, we design an “implicit keyframe learning with Taylor series expansion” mechanism that adaptively extracts key beat frames via neural networks and employs a third-order Taylor series for higher-order continuous motion synthesis; for spatial modeling, we adopt an anatomy-based partitioning strategy, dividing the human body into seven functional regions for independent and collaborative modeling to enhance complex pose-generation capabilities. Finally, a multiscale architecture integrates spatiotemporal features at different levels of granularity. Extensive experiments on the AIST++ and PopDanceSet datasets demonstrate that our method outperforms state-of-the-art approaches in the physical plausibility of generated motions (PFC and PBC), diversity (Div-k and Div-g), and the critical music–dance beat-alignment score (BAS), thereby validating the effectiveness and advantages of the proposed framework.