RSS Amplifier

Hornbill Dispatch by Farmers for Forests · Mar 27, 2026

Hornbill Dispatch #16: When Trees Started Talking Back to Us

0
Sign in to vote or save

Farmers for Forests · Hornbill Dispatch by Farmers for Forests

Last October, I got the opportunity to experience something pretty cool: attending Reid Hoffman's Masters of Scale Summit in San Francisco, listening to and meeting some of the foremost thinkers and doers of our times. And learning new terms like "zero-shot learning," and "AlphaFold moment.” I was there as one of 40 early-stage founders selected to attend, oscillating between taking notes and wondering if the person next to me could tell I’m furiously Google-ing every other word coming out of Fei-Fei Li’s mouth.

Left: Early Stage Founders Group Photo, Right: In conversation with fellow Early Stage Founder | Photo Credit: Masters of Scale

On the second day, Reid Hoffman interviewed Aza Raskin, Co-Founder of Earth Species Project - an organization using AI to decode how animals communicate. And Aza told us about a 1994 University of Hawaii study that made my brain short-circuit.

“Do something you’ve never done before. Together.”

That was the instruction researchers gave to dolphins.

Researchers had taught dolphins two gestures. The first gesture meant: “do something you’ve never done before” - to essentially “innovate.” The dolphins understood this. They could remember everything they'd done previously, understand the concept of negation, and invent new behaviors on command.

Pretty cool. But then the researchers taught them a second gesture: “do something together.”

And then they combined them: “Do something you’ve never done before. Together.”

Aza explained: “The dolphins would go down, exchange sonic information, come up and do the same thing that they had never done before at the same time.”

If dolphins are this sophisticated - coordinating innovation through sonic communication - then what else are we missing? How many other species have complex communication systems we simply can’t perceive or decode?

Aza Raskin at Masters of Scale Summit | Photo Credit: Masters of Scale

Turns out that’s the question driving the Earth Species Project (ESP). And Aza explained that answering it requires building fundamentally new AI tools. They’re developing large-scale foundation models - NatureLM - designed to work across the entire tree of life, bringing AI’s power to understanding non-human communication.

Wait. Foundation models? For animal communication? That’s... exactly what we’d been building. Just for trees instead of dolphins.

For the past four years, we’d been building tree detection AI. Training models on drone imagery. Creating ground truth datasets. Figuring out how to make models work across different species, different terrains, different lighting conditions.

I’d always thought of this as forestry work - restoration monitoring, carbon accounting, agroforestry optimization. But sitting there listening to Aza, I realized: we weren’t just solving forestry problems. We were solving the exact same AI problems that ESP was tackling for whale clicks and crow calls.

The technical challenges were identical. Here’s what Aza described about ESP’s approach - and how familiar each one sounded:

Aza told us about their work with crows at the University of León. Crows have remarkable social structures - collective child rearing like a commune, their own dialects, and they even take in outside adults and teach them the local dialect before those newcomers start participating in childcare.

“We work with the university to put backpacks on crows,” Aza explained. “And we get to record what they’re saying all the time.”

(Yes. Crow backpacks.)

Watch Earth Species Project's crow communication research in action here

“One of the things we’ve been finding is that around 70% of the calls that they [crows] make are actually quiet, intimate calls. When we brought these to the scientists we were working with, they said: ‘We haven’t seen these calls before.’”

Think about that. Scientists have been studying crows for decades. But they’d only been listening to the loud stuff - the alarm calls, the territorial announcements, the shouts. They were missing 70% of crow communication happening in quiet, close-range conversations.

“We’ve been trying to understand the language of another species by only listening to their shouts,” Aza said. “It’s like trying to understand humans by only listening to Twitter.”

We’d had the exact same problem with forests. For years, the restoration monitoring world has focused on what’s easy to detect from far away - satellite imagery showing large-scale deforestation, broad forest health metrics, regional carbon stock estimates. The “shouts.”

But what about the quiet signals? The subtle stress indicators in individual saplings? The microclimate variations that determine which trees thrive?

Satellites can’t see that granular detail. They miss significantly large parts of the story happening at the scale where restoration actually happened for us - on small farms, with individual trees, in the hands of farmers making daily decisions.

ESP put backpacks on crows to hear the quiet calls. We use cameras on drones for the same reason - to get close enough to see what satellite monitoring misses.

“The same tools we are building for crows, these methods can scale,” Aza explained. “The way AI works - learning about one species teaches you about all species. In human language, if I learn German and French and Japanese, I can learn very quickly Esperanto and Aramaic, because each language has a pattern. And the same thing is true of the animal kingdom.”

This is where the foundation model approach becomes powerful. ESP isn’t building a different AI for every animal. They’re building NatureLM - a model trained on diverse animal sounds that can transfer learning across species. The same architecture that detects beluga whale communication patterns can be adapted for crow calls, for elephant rumbles, for bat echolocation.

“We’re building these tools that are crossing the entire tree of life,” Aza continued. “And we’re heading into AlphaFold moments where we will be able to show that many species have names, or many species have compositional language, or many species have abstract thinking going from abstract to specific.”

(For the non-AI nerds: AlphaFold was the AI breakthrough that solved protein folding - a 50-year-old problem in biology. It was a watershed moment. Aza was predicting similar breakthroughs in understanding animal communication.)

We’d built the same approach for trees. We started with organized agroforestry plots in the tropical and subtropical zones of western Maharashtra - rows of sweet lime and mango saplings, predictable spacing, open canopy. Relatively simple detection scenarios. But the real test of a foundation model isn’t how well it works on training data that looks like itself. The test is: can it generalize to ecological contexts it’s never seen?

Because the world needs forest monitoring in places that don’t look like neat agroforestry rows. Dense old-growth sal forests in eastern Maharashtra where individual tree crowns blur into continuous canopy. Bamboo plantations in the Philippines where culms grow in tight clumps. Mixed species agroforestry in Tanzania where farmers are integrating indigenous trees with cash crops in patterns we’ve never encountered.

Old-growth forest in Korchi, Gadchiroli - the kind of dense canopy our models are now learning to read after training on organized agroforestry plots | F4F Field Photo

We can’t build a different detection model from scratch for every terrain type, every tree species, every altitude, every biome, every forest structure. The computational cost would be prohibitive - tens of thousands of new annotations per context, months of retraining, siloed models that can’t learn from each other’s improvements.

Instead, we’re piloting our Maharashtra-trained models in new regions - bamboo in the Philippines, mixed agroforestry in Kenya, dense forest in Gadchiroli - and using ground validation from local partners to understand where the model works well and where it needs fine-tuning. A few hundred validated examples from partners who know their landscapes intimately, rather than rebuilding everything from zero.

The foundation is transferable. Not because all forests look the same, but because the core computer vision task - detecting photosynthetically active biomass in overhead imagery, distinguishing individual plants from background, estimating structure from projections - follows learnable patterns that transcend specific contexts.

Just like ESP’s NatureLM can go from crow calls to whale songs, our detection models can go from agroforestry to old-growth forest. Different ecosystems, different structures, different challenges. Same foundational architecture, adapted through transfer learning rather than rebuilt from zero.

“Here’s an interesting puzzle,” Aza said. “Orcas can communicate up to 150 kHz. We can hear up to 20 kHz. But in water, high frequencies don’t go very far. So why are they communicating at such a high pitch?”

The theory: because they’re often communicating in close calls. Intimate conversations, while moving around each other, and often touching. High frequency means short range, which means private. They’re not broadcasting - they’re texting.

But this creates the “cocktail party problem” - when you record whales in the wild, you don’t get clean, isolated vocalizations. You get chaos. Multiple animals calling at once, at different frequencies, overlapping. ESP had to build AI that could separate individual animal voices from this acoustic mess.

We had the exact same problem - just visual instead of acoustic.

When you fly a drone over an agroforestry plot, you don’t get neat rows of evenly-spaced trees. You get chaos. Trees overlapping with each other. Shadows creating false signals. Farmers intercropping saplings with vegetables - brinjal with their characteristic beautiful purple flowers emerging from fields of young papaya trees (see video below from a F4F farmer plot), sweet lime between rows of soybean, custard apple surrounded by pigeon pea. Varying canopy densities. Different heights creating occlusions where larger trees hide smaller ones.

And then there’s the camouflage problem: young saplings get visually lost in the spectral signature of whatever’s growing around them. Noxious weeds like Parthenium hysterophorus grow aggressively in the monsoon, completely obscuring saplings. Cover crops that farmers intentionally plant for soil health look remarkably similar to young trees in RGB imagery. Ground vegetation responding to the same irrigation that’s feeding the trees creates a uniform green background that makes crown detection nearly impossible without careful analysis of leaf texture, growth patterns, and spectral signatures.

Photo of Parthenium hysterophorus (commonly known as Congress grass) overtaking one of F4F plots

Our model learned to separate individual tree crowns from dense, overlapping canopies. To distinguish between a tree and a tree’s shadow. To detect small saplings next to large mature trees without the large tree’s canopy drowning out the small one’s signature.

Different modality - sound waves vs. pixels. Identical challenge - separating individual signals from noisy, overlapping data.

ESP needs to decode communication across thousands of species - whales, crows, elephants, dolphins, birds, bats, primates. Each with different vocal structures, different contexts, different behavioral repertoires. They can’t build proprietary models for every species and lock them behind paywalls. The problem is too big.

We need to monitor forests across every biome on Earth - tropical, temperate, boreal, arid. Across every terrain type, every tree species, every stage of restoration. We can’t build proprietary detection systems for every context and charge licensing fees. The problem is too big.

The only way solutions to these problems scale is if everyone can build on everyone else’s work. If a researcher in Tanzania can take ESP’s crow model and adapt it for local bird species. If an implementer in the Philippines can take our Maharashtra model and fine-tune it for bamboo. If improvements flow back into the commons, making the foundation stronger for everyone.

That’s why ESP is making everything open-source - benchmark datasets like BEANS (Benchmark for Animal Sounds) that allow researchers to standardize evaluation methods, foundation models like AVES (Animal Vocalization Encoder based on Self-Supervision) that work across species, and the acoustic datasets they’ve collected.

And that’s why we had independently arrived at the exact same conclusion early last year. Not because it’s idealistic. Because it’s the only approach that works at the scale these problems demand.

…currently with Project Anāhitā:

Our white paper is published and freely available - the complete technical methodology behind AI-based carbon stock calculation in agroforestry plantations, including our validation results and honest discussions of limitations.

Our ground truth dataset - 7,033 manually measured trees across Maharashtra - is open-sourced on GitHub. GPS coordinates, species identification across 23 species, field-verified DBH measurements, age estimates. Hundreds of person-hours in 45°C heat, freely available to anyone who wants to build on it.

Our complete codebase - the tree detection models, the Gaussian Process implementations, the data processing pipelines - is coming. We’re documenting it properly, which is taking longer than we’d like. But it’s coming, with documentation that we genuinely hope is better than the usual “it works on my machine” approach.

And today, we're releasing TreeLens - free for anyone to use. More on that below.

For the past four instalments of this series, we’ve been telling you about our technical journey:

Today, we’re releasing the platform that makes all of that accessible - to anyone with a drone and a plantation to monitor.

TreeLens is simple to use. Here’s how it works: you fly a drone over your plantation, create an orthomosaic from the images (free tools like WebODM or a paid one like Pix4D can do this), and upload it to TreeLens at farmersforforests.com/tree-lens. The platform accepts .tiff, .png, and .jpg files up to 100 MB, and works best with orthomosaics at around 2.5 cm resolution - the kind any standard agricultural drone can produce. If you also have a Digital Surface Model (DSM) file for your plot, upload that too: it’s what gives us accurate tree height, and from height, biomass and carbon estimates.

TreeLens in action

What you get back: individual tree detections and counts, species identification across 9 species, height and crown size for each detected tree, and carbon sequestration estimates.

Right now, our models are optimised for managed agroforestry plantations - the organised plots we’ve been training on for four years. Natural forest detection is under active development.

A note on access: TreeLens is free to use - anyone can create an account at farmersforforests.com/tree-lens. Currently, each account supports up to 3 file uploads (max 100 MB each), which is enough to run a meaningful analysis on your plots. You can test it immediately using our pre-loaded drone images, or upload your own orthomosaic. For any questions or concerns, please reach out at tech@farmersforforests.com.

At an event a few weeks ago, I met someone from Pula - an organisation making crop insurance work for smallholder farmers across Africa and Asia.

Their insight was deceptively simple. The reason smallholder farmers couldn’t access insurance wasn’t that they were uninsurable. It was that nobody had the data to price their risk. Traditional insurers need to know what’s on the ground - what was planted, how it’s growing, what a normal season looks like - before they can underwrite anything. Pula’s solution was to generate that data and connect it to insurers who could now, for the first time, see the risk clearly enough to price it.

Nature-based projects face a version of the same problem. The trees are real. The carbon is real. The farmer’s commitment is real. But from the perspective of a lender, an insurer, or a government scheme trying to direct subsidies efficiently - none of it is legible. You can’t collateralise what you can’t verify. You can’t underwrite what you can’t measure. This is why, despite enormous enthusiasm for nature-based climate solutions, a lot of financing that reaches farmers still comes from philanthropy - not because commercial capital doesn’t care about trees, but because the verification layer has been missing.

TreeLens is an attempt to build that missing layer. Consider a smallholder agroforestry farmer in Vidarbha with 400 trees on her land, three years into a five-year project. Today, she is largely invisible to formal financial systems. But with verified, timestamped drone data showing exactly which trees are alive, how tall they are, how much carbon they’ve sequestered - she suddenly has something that looks, to a bank or insurer, like an asset with a track record. Maybe better terms on a crop loan. Maybe parametric insurance tied to tree survival rates, not just rainfall. Maybe a government scheme that can verify outcomes rather than just inputs. The same detection layer that produces a carbon estimate can underpin a microinsurance product, a green loan. The asset is the same. What changes is who can now see it.

We think of this as Detect, Debt, Deploy. Detection creates the verified data layer. That layer makes nature-based projects legible to capital - reducing perceived risk enough that commercial and blended finance can flow in alongside philanthropy. That capital then deploys at the scale philanthropy alone will never reach.

None of this is solved yet. The financial instruments are nascent, the integration between monitoring technology and financial services is still being built, and the last-mile challenges of getting capital to smallholder farmers are every bit as hard as the technical ones. But the detection layer - the piece that has historically been missing - now exists.

When I was listening to Aza Raskin talk about dolphins, what struck me most was the instruction the researchers gave: Do something you’ve never done before. Together.

The dolphins understood something that a lot of organisations working on climate and restoration still struggle with: that the most interesting things - the breakthroughs, the new behaviours, the solutions nobody has tried yet - only emerge when you share information and act in coordination. The dolphins went underwater, exchanged sonic information, and surfaced doing the same new thing at the same time. Coordinated novelty. That’s what open-source is, for us.

We spent four years building this, funded by patient philanthropists who believed in a possibility rather than a product. We didn’t build it to keep it. We built it because the problem - smallholder farmers excluded from the very climate solutions their land stewardship makes possible - is too large and too urgent for any one organisation to solve.

Project Anāhitā is now live in pieces: the white paper is published, the dataset is on GitHub, the codebase is coming, and TreeLens is open to partners who want to try it. We’re not done. We’re probably not even halfway. But we’ve learned enough to know that the next phase of this work won’t be built by us alone.

If you’re an implementer working with smallholder farmers and want to try TreeLens on your plots - reach out. If you’re a researcher who sees something interesting in our methodology or our dataset - build on it, and tell us what you find. If you’re working on the financing side - the insurance products, the green loans, the blended finance vehicles - and you’re trying to figure out what verified tree data could unlock for you, we’d genuinely love that conversation.

And if you've made it to the end of a five-part series about drone orthomosaics and Gaussian Process Regression - honestly, we're a little worried about you. But also, thank you.

- Arti Dhar, Co-Founder

No posts

Read the original on farmersforforests.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.