A San Francisco Bay Area startup that promises to solve one of generative video's thorniest problems — soundtracks that actually match what's on screen — has closed an $11 million seed round led by B Capital, with backing from Redpoint. Sonilo declined to disclose its valuation.
The pitch is straightforward enough: Upload footage, and the company's models analyze cuts, motion and pacing to generate music or sound effects that land precisely on the final frame. It's the kind of technical problem that sounds mundane until you've watched AI-generated videos with soundtracks that drift awkwardly out of sync or fail to register a scene change.
Sonilo launched its v1.0 model on March 15, 2026. The company now offers video-to-music and text-to-music generation, plus a video-to-sound-effects pipeline that launched on July 25, 2026.
Computing Power and Licensing Deals
The fresh capital will fund the computing infrastructure needed to scale the models, according to an October 8 announcement. Sonilo also plans to expand across creator and developer platforms while building out music licensing arrangements.
"We built Sonilo to understand the visual story," CEO Shawn Song told Variety. "You give us the footage, and Sonilo understands what's happening frame by frame."
That frame-level analysis is the technical hook, though it remains to be seen whether creators will embrace yet another AI tool in an increasingly crowded market for generative audio.
The TikTok Connection

Song's background includes leading multimodal AI work at TikTok, according to Variety, and training at Carnegie Mellon. His co-founders bring complementary experience. Chief Operating Officer Keli Li built and ran TikTok's global music operation. Chief Technology Officer Alex (Zongyu) Yin holds a PhD in Computer Music from the University of York, which he completed between 2018 and 2021; he won the school's Best PhD Thesis and KM Stott Prize in 2023, per his LinkedIn profile. Chief Marketing Officer Trista Taylor worked on consumer AI applications that landed on a16z's Top 50 rankings.
The company lists between 11 and 50 employees on LinkedIn and operates from Menlo Park, California. Sonilo was founded in 2026.
The Copyright Angle

Sonilo partnered with Shutterstock in May to license the stock platform's music catalog for model training — Shutterstock's first such deal with a video-to-music AI company, according to a company blog post. The startup positions itself as "built for licensed commercial use from day one," emphasizing rights-holder participation and revenue sharing on its website.
That's a pointed contrast with text-to-music competitors Suno and Udio, both facing ongoing copyright lawsuits from major labels. Suno raised more than $400 million at a $5.4 billion valuation in June, Reuters reported, even as Universal and Sony sued the company for a second time in September. The labels claimed Suno's v6 models "are the fruit of the same poisoned tree," according to Music Business Worldwide.
Whether licensing deals will insulate Sonilo from similar legal challenges is an open question. The music industry has shown little patience for AI companies training on copyrighted material, licensed or otherwise, when those models can generate output that competes with existing catalogs.
"Sonilo stands out because its multimodal technology gives creators two ways to generate music and sound effects," said Daisy Cai, general partner at B Capital, in the Variety report.
Distribution Through Developers

Sonilo released a public API in May with a REST endpoint for video-to-music generation. Parameters include automatic ducking under dialogue, according to developer documentation. The platform added a native ComfyUI node around the same time and integrated with fal.ai in June, with fal hosting Sonilo's model on its generative media infrastructure.
In August, Scenario added Sonilo's v1.1 text-to-music and video-to-music models to its platform, according to a LinkedIn post. The company also lists plug-ins or integrations for Adobe Premiere Pro, Unity, Godot, Epic Games Fab, Roblox and Defold on its homepage, though the depth and adoption of those integrations is difficult to verify from public information.
Product updates arrived quickly between September and October: batch processing for up to 10 videos, section-based scoring, dubbing into 23 languages with lip sync, and expanded upload limits to 30 minutes on premium plans, according to the company's news archive. Consumer pricing starts at $11.99 per month.
Song presented at the Venice Film Festival's AI Forum in September, where Sonilo provided the soundtrack for a TapNow trailer. "As video generation becomes faster and more automated, the need for real-time, adaptive music has become increasingly critical," he said in an earlier product launch statement.
The company is betting that need will translate into sustainable revenue before the computing costs of scaling generative models eat through its seed funding. It's a familiar calculus for AI startups in 2024, and one that has yet to resolve for most players in the space.
