I spent months trying to identify a bird in my backyard. I built an Alexa skill so I’d never have to do it again.
A pair of green birds showed up every morning on the same wire, near my house. I was sure they weren’t maritacas (plain parakeets) — the size was different, the noise was different. But I didn’t know their name.
It took months. I searched WikiAves, which is the reference in Brazil and excellent work — but for a beginner like me, the sheer volume of information was more intimidating than useful. I didn’t know the right terms to filter by. I didn’t know whether “green with a red patch on the wing” was an exotic sabiá (thrush) or a common parrot. I found out later: they were Maracanã parakeets.
And then came the irony that planted the whole thing: I went to check what their call sounded like so I could confirm it, and found out that parrots don’t have a standardized call. The song I wanted to use as proof simply doesn’t exist in any classifiable form. The challenge was already in my head.
The first idea was wrong (and unethical)
My original plan was to use Alexa to listen to the bird sounds around me and identify them by audio. Two problems killed that idea the same day:
Alexa only picks up human voice. It doesn’t listen to the environment — the microphone is optimized for spoken commands. Background sounds are noise, not signal. There is no “listening mode” I could turn on.
Attracting birds with playback is a questionable practice. Playing a species’ song to attract it causes territorial stress, can drive birds away from their nests and interferes with breeding. IN ICMBio 14/2018 (a regulation from Brazil’s federal biodiversity agency) and CEMAVE (its bird research and conservation center) advise against it. When I found that out, I dropped the idea without a second thought.
So the skill became something else: instead of “record what you heard”, it became “describe what you saw”. You talk to Alexa in natural language — color, size, where it was, bill shape — and it returns the most likely species.
The stack: Python, Bayes and four open databases that don’t talk to each other
Alexa isn’t exactly a smart device. It recognizes voice and calls a function in the cloud. All the intelligence sits in a Lambda (Python) on AWS, and the real work was making four databases that were never built to talk to each other produce a coherent answer.
AVONET (Tobias et al. 2022, CC BY 4.0) has standardized morphological traits for more than 11,000 species — that’s where the 7 attributes the skill extracts from your description come from: primary color, secondary color, size by anchor, habitat, posture, bill and tail.
GBIF has occurrence records for birds in Brazil. I used that data as a Bayesian prior: if you’re in São Paulo and describe “a big green bird”, the scorer gives more weight to the species that actually show up in the region (30 km radius, with at least 20 records). Bem-te-vi (great kiskadee) before a rare species with the same description.
xeno-canto has the real song recordings. And here comes a decision that wasn’t optional for me: I couldn’t just hook into the audio without crediting whoever recorded it. Every playback shows the recordist’s name, the license (CC BY-NC-SA 4.0) and the recording number in the archive. Python turned out to be great for pulling those audio files and organizing both species and authors programmatically.
Han et al. 2025 (CC0) fills in with plumage color data that AVONET doesn’t cover.
The AMAZON.SearchQuery slot type in pt-BR captures free speech — “I saw a yellow little bird with black wings on the ground in the backyard” comes in whole as text and the parser extracts the attributes. There’s no menu, no “say 1 for color”. You speak the way you’d speak to a friend.
The numbers, with nothing made up
1,627 species from all of Brazil. 1,014 with a real xeno-canto recording. The database started with 390 species from the São Paulo region and was expanded to cover every Brazilian biome.
Accuracy simulated with realistic noise (incomplete descriptions, wrong attributes, regional synonyms): the right bird shows up in the top 3 suggestions in ~54% of cases. It’s not 90%. It’s what the model delivers today with open data and a parser that accepts free speech. Being transparent about that number matters more than rounding it up.
There’s also a song quiz: the skill plays a real recording and you try to guess the species. Get it right and it keeps score. Get it wrong and it tells you what it was and plays it again so you can make the association. It’s a way to train your ear without having to be out in the field.
And the biggest engineering bottleneck wasn’t the scorer — it was the Alexa Developer Console. Testing an Alexa skill is a slow process, with deploys that drag, logs that lag and an interface few people have mastered. Anyone who has developed for Alexa knows. Anyone who hasn’t will find out.
What the skill does NOT do (on purpose)
It doesn’t record the bird’s sound — Alexa can’t. It doesn’t loop songs to attract birds — that’s unethical. It has no usage metrics because it was just published. It has no user testimonials because I’m not going to make them up. WikiAves, xeno-canto, GBIF and AVONET are credited sources, not partners.
Every song playback comes with a responsible-playback notice. Single playback, controlled volume, credit to the recordist, guidance about the breeding season. That isn’t a feature — it’s an obligation.
The invitation: test it, break it, improve it
The code is MIT and the repository is public: github.com/jvitorcarvalho/aves-do-brasil-alexa.
I want people to have fun with this. To discover the call of the urutau (the common potoo), which is one of the most curious things you’ll ever hear. To identify a bird in the backyard and tell a friend over lunch. For a developer in Manaus to grab the repository and calibrate the priors for the Amazon. For an ornithologist to fix the common names that came from GBIF and probably have errors.
Today the skill covers all of Brazil — 1,627 species, 1,014 with a recording. There’s room.
To try it: “Alexa, abre Aves do Brasil”. It’s free, no ads, no data collection.
To contribute: open an issue, send a PR, or tell me that the common name for the tico-tico (rufous-collared sparrow) in your town is something else. It all helps.
Developed by Evolutiva Negócios Digitais (João Monlevade/MG, Brazil). Sources: AVONET (Tobias et al. 2022, CC BY 4.0), GBIF (CC0), xeno-canto (CC BY-NC-SA 4.0), Han et al. 2025 (CC0), WikiAves (public search). Playback notice follows IN ICMBio 14/2018 and CEMAVE guidelines.
Frequently asked questions
How does the Aves do Brasil skill identify a bird from a spoken description?
The skill uses a Bayesian scorer that cross-references 7 attributes extracted from your speech — color, size, habitat, posture, bill, tail and secondary color — with morphological data for 1,627 Brazilian species. The data comes from AVONET (Tobias et al. 2022), complemented by GBIF occurrence data to weight the species most likely in your region. You describe what you saw in natural language and get back the species that best match the description.
Does the skill play songs to attract birds?
No. The skill plays real xeno-canto recordings only once per query, with controlled volume and credit to the original recordist. Looping songs to attract birds causes territorial stress and can interfere with breeding — a practice discouraged by IN ICMBio 14/2018 and by CEMAVE. Every playback comes with a responsible-playback notice.
What is the hit rate of the voice-based identification?
In simulations with realistic noise (incomplete descriptions, wrong attributes, regional synonyms), the right bird shows up in the top 3 suggestions in roughly 54% of cases. That number reflects the current performance with open data and a free-speech parser. The skill doesn’t round its metrics up — being transparent about the limitations is intentional.
Is the Aves do Brasil skill free, and does it collect personal data?
The skill is free, with no ads and no personal data collection. The code is open source under the MIT license, available in the public repository. The databases used (AVONET, GBIF, xeno-canto, Han et al. 2025) are all open and properly credited.
Can I contribute to the skill or correct common bird names?
Yes. The repository accepts issues and pull requests — contributions from ornithologists, developers and birdwatchers are welcome. Regional common names that differ from the GBIF records can be corrected directly. The current database covers 1,627 species from Brazil and 1,014 with a real recording, but there’s room for regional calibration and expansion.
