Kalpa Labs is an audio research lab working towards generalist audio models: models that follow instructions and learn new tasks in context. Our first conversational speech models are in beta today. Direct a scene in Studio, talk to a model live, or build on the API.
Most voice cloning products only preserve speaker identity, and fail to mimic speaker delivery nuances. To stress test our voice cloning capabilities, we intentionally clone "meme voices" that say things in an exaggerated or distorted manner for a memetic effect.
Speech models today are roughly where language models were a few years ago: they can speak, but they can't really listen. They don't follow instructions about sound, learn a voice or a style from a few examples in context, or notice how something was said and not just what. Generalist audio models that close that gap is our whole mission. Read more →
Until then, the beta is open. Go talk to it→