Behavior-Conditioned Animal Vocalization Generation: Event-Aware Modeling and Multidimensional Evaluation
Generating short, duration-variable animal vocalizations from natural-language descriptions of behaviors, social interactions, and contexts.
- 01BeVo dataset
- 02Event-Aware generation
- 03Multidimensional evaluation
01 / LISTEN
Listen by behavior
Explore 42 condition labels across six species. Choose a label to compare a real test recording with its corresponding final Event-Aware model output.
Condition label
Choose a categoryGeneration prompt · exact test manifest text
A vocalization is heard from a spotted hyena when it seeks distant clan contact.
- Species
- Spotted hyena
- Anchor
- Whoop
- Real duration
- —
Each pair shares playback gain, STFT settings, axes, and an 80 dB color scale with a common amplitude reference.
02 / OVERVIEW
From field annotations to conditioned generation
Behavior-conditioned animal vocalization generation aims to synthesize vocalizations from natural-language descriptions of animal behaviors, social interactions, and contexts.
We establish a pipeline spanning data construction, generation, and evaluation. Public bioacoustic data from six species is linked with species-specific ethological and bioacoustic literature to form natural-language behavioral conditions. A rectified-flow generator is then adapted with Event-Aware variable-length modeling for short, duration-variable vocalizations.
BeVo dataset
Literature-grounded behavioral conditions linked to public bioacoustic recordings.
Event-Aware modeling
Actual-duration encoding with masked padded latent positions.
Multidimensional evaluation
Discriminability, acoustic fidelity, and language-aligned representation fidelity.
03 / DATASET
The BeVo Dataset
A standardized collection of vocalization events connected to literature-grounded behavior cards through 42 species-specific vocalization anchors.
Spotted hyena
12,242 events
Meerkat
5,639 events
Common marmoset
6,000 events
Goat
3,586 events
Plains zebra
615 events
Zebra finch
3,109 events
Table 1 · Dataset splits24,968 train · 3,123 validation · 3,100 test
| Split | Goat | Hyena | Marmoset | Meerkat | Zebra | Zebra Finch | Total |
|---|---|---|---|---|---|---|---|
| Train | 2,826 | 10,216 | 4,800 | 4,475 | 353 | 2,298 | 24,968 |
| Validation | 360 | 1,226 | 600 | 564 | 62 | 311 | 3,123 |
| Test | 400 | 800 | 600 | 600 | 200 | 500 | 3,100 |
| Total | 3,586 | 12,242 | 6,000 | 5,639 | 615 | 3,109 | 31,191 |
Samples are split within each species and vocalization anchor using five duration-quantile strata.
Hyena and meerkat recordings use annotated temporal boundaries; the other corpora provide pre-segmented or short vocalization recordings. Curation includes species- and anchor-specific RMS screening, AnimalCLAP or BioLingual filtering where suitable, and downsampling. Exact preprocessing thresholds are not yet published.
04 / METHOD
Event-Aware variable-length modeling
Short animal vocalizations are encoded at their actual durations so that padded positions neither act as valid context nor contribute to the rectified-flow objective.
- 01
Encode actual duration
Each vocalization is independently encoded with only the padding needed for VAE hop alignment.
- 02
Pad within the minibatch
Variable-length latent sequences are padded only to the longest event in the current batch.
- 03
Mask attention and loss
Padding is excluded from transformer context and rectified-flow optimization.
- 04
Generate for the requested duration
Inference initializes the latent sequence for the target event duration.
Experimental setup
- Training
- 400 epochs · batch size 64
- Optimizer
- AdamW · learning rate 1 × 10−5
- Generation
- 50 steps · guidance scale 4.5
- Target duration
- Corresponding real test recording
The VAE and FLAN-T5 text encoder are frozen; the Flux Transformer, text projection layer, and duration embedder are trained. Both adapted systems share training-sample exposure and condition sampling. The checkpoint with the lowest validation macro loss is selected. Each test instance uses one fixed prompt–duration pair across systems.
05 / RESULTS
Results at a glance
Task-specific adaptation improves all evaluated metrics over zero-shot generation. Event-Aware modeling further improves vocalization discriminability and acoustic fidelity over standard fine-tuning, while language-aligned fidelity metrics remain nearly unchanged.
Anchor macro accuracy ↑
61.99%59.86% fine-tunedFAD ↓
1.22391.4047 fine-tunedAFDD ↓
0.92061.2480 fine-tunedSwap increase rate
85.39%Cases with increased new-target probabilityAVEX Δptgt +0.3139| System | Macro Acc. ↑ | mAP ↑ | FAD ↓ | AFDD ↓ | AnimalCLAP-FAD ↓ | BioLingual-FAD ↓ |
|---|---|---|---|---|---|---|
| Real Reference | 0.8852 | 0.9240 | 0.2321 | 0.2266 | 0.0075 | 0.0275 |
| TangoFlux Zero-shot | 0.1787 | 0.2022 | 6.7080 | 3.4759 | 0.3318 | 0.4904 |
| TangoFlux Fine-tuned | 0.5986 | 0.6345 | 1.4047 | 1.2480 | 0.0491 | 0.1237 |
| TangoFlux Event-Aware | 0.6199 | 0.6551 | 1.2239 | 0.9206 | 0.0492 | 0.1231 |
Bold marks the best generated-system result. Real Reference distribution metrics compare real test audio with an equally sized training subset. Scroll horizontally to view every metric.
Table 3 · Species-wise performanceReal Reference / Event-Aware
| Species | Macro Acc. ↑ | FAD ↓ | BioLingual-FAD ↓ |
|---|---|---|---|
| Hyena | 0.9000 / 0.5500 | 0.1622 / 1.1566 | 0.0227 / 0.1029 |
| Meerkat | 0.7883 / 0.6967 | 0.2478 / 1.5149 | 0.0155 / 0.0739 |
| Marmoset | 0.9717 / 0.6550 | 0.1147 / 1.8304 | 0.0225 / 0.2129 |
| Goat | 0.8300 / 0.4050 | 0.1561 / 0.7219 | 0.0273 / 0.0974 |
| Zebra | 0.9650 / 0.6550 | 0.5648 / 1.1341 | 0.0520 / 0.1151 |
| Zebra Finch | 0.8560 / 0.7580 | 0.1470 / 0.9853 | 0.0249 / 0.1364 |
Generation difficulty varies across species and evaluation dimensions; each cell is Real Reference / Event-Aware.
Representation analysis
Real and generated vocalizations share broad species-level organization.
A joint t-SNE visualization of AnimalCLAP embeddings uses circles for real samples and triangles for generated samples.
06 / SCOPE
What this demo shows
Associated conditions
Behavioral conditions describe groups of recordings sharing a vocalization anchor, not verified behavior for each event.
Linguistic generalization
Evaluation uses new formulations of behavioral knowledge represented during training, not entirely unseen behaviors.
Distribution-level fidelity
CLAP-FAD measures species-wise distribution fidelity in language-aligned spaces, not sample-wise prompt correctness.
07 / CITATION
Cite this work
Use this manuscript citation. Venue details can be added when available.
@article{deng2026bevo,
title = {Behavior-Conditioned Animal Vocalization Generation: Event-Aware Modeling and Multidimensional Evaluation},
author = {Deng, Xinlong and Tian, Yong and Guan, Jian and Kong, Qiuqiang and Jiang, Jie and Cao, Yin},
year = {2026},
note = {Manuscript},
url = {https://persimmontian.github.io/BCAVG/}
}