Research project · Bioacoustics

Project page / 2026

Behavior-Conditioned Animal Vocalization Generation: Event-Aware Modeling and Multidimensional Evaluation

Generating short, duration-variable animal vocalizations from natural-language descriptions of behaviors, social interactions, and contexts.

  1. 01BeVo dataset
  2. 02Event-Aware generation
  3. 03Multidimensional evaluation
Hyena Meerkat Marmoset Goat Zebra Zebra Finch
6species
31,191vocalization events
42vocalization anchors
524.3minutes of audio

01 / LISTEN

Listen by behavior

Explore 42 condition labels across six species. Choose a label to compare a real test recording with its corresponding final Event-Aware model output.

Condition label

Choose a category

Generation prompt · exact test manifest text

A vocalization is heard from a spotted hyena when it seeks distant clan contact.

Species
Spotted hyena
Anchor
Whoop
Real duration
—
A / REAL AUDIO

Real vocalization

Listen now

B / GENERATED AUDIO

Final Event-Aware model

Listen now

02 / OVERVIEW

From field annotations to conditioned generation

Behavior-conditioned animal vocalization generation aims to synthesize vocalizations from natural-language descriptions of animal behaviors, social interactions, and contexts.

We establish a pipeline spanning data construction, generation, and evaluation. Public bioacoustic data from six species is linked with species-specific ethological and bioacoustic literature to form natural-language behavioral conditions. A rectified-flow generator is then adapted with Event-Aware variable-length modeling for short, duration-variable vocalizations.

01

BeVo dataset

Literature-grounded behavioral conditions linked to public bioacoustic recordings.

02

Event-Aware modeling

Actual-duration encoding with masked padded latent positions.

03

Multidimensional evaluation

Discriminability, acoustic fidelity, and language-aligned representation fidelity.

03 / DATASET

The BeVo Dataset

A standardized collection of vocalization events connected to literature-grounded behavior cards through 42 species-specific vocalization anchors.

Overview of the BeVo dataset construction pipeline
Figure 1 Overview of the BeVo Dataset.
01

Spotted hyena

12,242 events

02

Meerkat

5,639 events

03

Common marmoset

6,000 events

04

Goat

3,586 events

05

Plains zebra

615 events

06

Zebra finch

3,109 events

Table 1 · Dataset splits24,968 train · 3,123 validation · 3,100 test
Number of vocalization events in each data split.
SplitGoatHyenaMarmosetMeerkatZebraZebra FinchTotal
Train2,82610,2164,8004,4753532,29824,968
Validation3601,226600564623113,123
Test4008006006002005003,100
Total3,58612,2426,0005,6396153,10931,191

Samples are split within each species and vocalization anchor using five duration-quantile strata.

Hyena and meerkat recordings use annotated temporal boundaries; the other corpora provide pre-segmented or short vocalization recordings. Curation includes species- and anchor-specific RMS screening, AnimalCLAP or BioLingual filtering where suitable, and downsampling. Exact preprocessing thresholds are not yet published.

04 / METHOD

Event-Aware variable-length modeling

Short animal vocalizations are encoded at their actual durations so that padded positions neither act as valid context nor contribute to the rectified-flow objective.

Event-Aware latent modeling pipeline
Figure 2 Actual-duration waveform encoding and masked latent flow.
  1. 01

    Encode actual duration

    Each vocalization is independently encoded with only the padding needed for VAE hop alignment.

  2. 02

    Pad within the minibatch

    Variable-length latent sequences are padded only to the longest event in the current batch.

  3. 03

    Mask attention and loss

    Padding is excluded from transformer context and rectified-flow optimization.

  4. 04

    Generate for the requested duration

    Inference initializes the latent sequence for the target event duration.

Experimental setup

Training
400 epochs · batch size 64
Optimizer
AdamW · learning rate 1 × 10−5
Generation
50 steps · guidance scale 4.5
Target duration
Corresponding real test recording

The VAE and FLAN-T5 text encoder are frozen; the Flux Transformer, text projection layer, and duration embedder are trained. Both adapted systems share training-sample exposure and condition sampling. The checkpoint with the lowest validation macro loss is selected. Each test instance uses one fixed prompt–duration pair across systems.

05 / RESULTS

Results at a glance

Task-specific adaptation improves all evaluated metrics over zero-shot generation. Event-Aware modeling further improves vocalization discriminability and acoustic fidelity over standard fine-tuning, while language-aligned fidelity metrics remain nearly unchanged.

Anchor macro accuracy ↑

61.99%59.86% fine-tuned

FAD ↓

1.22391.4047 fine-tuned

AFDD ↓

0.92061.2480 fine-tuned

Swap increase rate

85.39%Cases with increased new-target probabilityAVEX Δptgt +0.3139
Table 2 · Overall evaluation↑ higher is better · ↓ lower is better
Overall results on behavior-conditioned animal vocalization generation.
SystemMacro Acc. ↑mAP ↑FAD ↓AFDD ↓AnimalCLAP-FAD ↓BioLingual-FAD ↓
Real Reference0.88520.92400.23210.22660.00750.0275
TangoFlux Zero-shot0.17870.20226.70803.47590.33180.4904
TangoFlux Fine-tuned0.59860.63451.40471.24800.04910.1237
TangoFlux Event-Aware0.61990.65511.22390.92060.04920.1231

Bold marks the best generated-system result. Real Reference distribution metrics compare real test audio with an equally sized training subset. Scroll horizontally to view every metric.

Table 3 · Species-wise performanceReal Reference / Event-Aware
Species-wise performance of TangoFlux Event-Aware. Each cell reports Real Reference / Event-Aware.
SpeciesMacro Acc. ↑FAD ↓BioLingual-FAD ↓
Hyena0.9000 / 0.55000.1622 / 1.15660.0227 / 0.1029
Meerkat0.7883 / 0.69670.2478 / 1.51490.0155 / 0.0739
Marmoset0.9717 / 0.65500.1147 / 1.83040.0225 / 0.2129
Goat0.8300 / 0.40500.1561 / 0.72190.0273 / 0.0974
Zebra0.9650 / 0.65500.5648 / 1.13410.0520 / 0.1151
Zebra Finch0.8560 / 0.75800.1470 / 0.98530.0249 / 0.1364

Generation difficulty varies across species and evaluation dimensions; each cell is Real Reference / Event-Aware.

Representation analysis

Real and generated vocalizations share broad species-level organization.

A joint t-SNE visualization of AnimalCLAP embeddings uses circles for real samples and triangles for generated samples.

t-SNE of real and generated vocalizations in AnimalCLAP space
Figure 3 AnimalCLAP representation space.

06 / SCOPE

What this demo shows

01

Associated conditions

Behavioral conditions describe groups of recordings sharing a vocalization anchor, not verified behavior for each event.

02

Linguistic generalization

Evaluation uses new formulations of behavioral knowledge represented during training, not entirely unseen behaviors.

03

Distribution-level fidelity

CLAP-FAD measures species-wise distribution fidelity in language-aligned spaces, not sample-wise prompt correctness.

07 / CITATION

Cite this work

Use this manuscript citation. Venue details can be added when available.

@article{deng2026bevo,
  title   = {Behavior-Conditioned Animal Vocalization Generation: Event-Aware Modeling and Multidimensional Evaluation},
  author  = {Deng, Xinlong and Tian, Yong and Guan, Jian and Kong, Qiuqiang and Jiang, Jie and Cao, Yin},
  year    = {2026},
  note    = {Manuscript},
  url     = {https://persimmontian.github.io/BCAVG/}
}