When we last left you, we'd successfully migrated from DeepForest to Detectron2, solved the "gift-wrapping a garden hose" problem of drip irrigation detection, and could finally count trees with 92% accuracy.
We were feeling pretty smug about ourselves. Our AI could spot trees, measure their heights, and even identify a few species. We'd gone from muddy QR codes to sophisticated computer vision in just two years. Time to pop the champagne, right?
Wrong.
We still weren’t sure how to measure the carbon in our trees. Here’s why.
Picture this: You're trying to estimate how much money is in someone's wallet by looking at them from across the street. You can see their height, you can guess their age. You might even know what kind of job they have(?). But the actual cash? That's hidden in their back pocket.
This was exactly our problem with trees.
(Yes, we're comparing trees to people's wallets. And yes, we couldn’t come up with a smarter analogy. Stay with us though!)
The amount of carbon in a tree is a function of the size of the tree trunk and the density of the wood. And so, to measure the carbon accurately we would need to know the diameter of the tree trunk and also know what species of tree it is, so that we can factor in the wood density.
Our drones flying overhead could see everything except the most important measurement for calculating carbon storage: the trunk diameter. It's literally hidden beneath the canopy, invisible to our aerial eyes. And without trunk diameter (or DBH - Diameter at Breast Height, as forestry folks call it), calculating how much carbon a tree stores without knowing the DBH is like trying to bake a cake without knowing how much flour to use in the recipe (lame analogy #2).
The run-of-the-mill solution would be to re-send our field team to these 55,000 trees we had annotated (to train the algorithm to detect trees) - but this time with measuring tapes to wrap around every single trunk. And that too in the horrid Maharashtra heat.
But instead, we asked ourselves what if we could teach our models to make predictions from visible data? Really, really sophisticated predictions.
In early 2024, we realized we needed someone who spoke fluent "uncertainty." Someone who could teach our models not just to predict trunk diameter from crown size and height, but to also say "I'm 85% confident in this prediction" or "This tree is weird, I'm only 60% sure."
Enter Dr. Abhishek Gupta from the Indian Institute of Technology (IIT) Goa, our mathematician who speaks fluent "uncertainty."
The breakthrough came in a serendipitous brainstorming session between Abhishek and Makarand Datar, our VP of Wildlife & Tech. As they discussed where uncertainty quantification would be most useful in our pipeline, old academic memories started surfacing. Makarand's master's thesis had involved something about Gaussian Processes and vehicle dynamics - dusty knowledge from years past that suddenly seemed relevant to trees.
The conversation evolved, ideas bounced back and forth, and eventually they landed on applying Gaussian Process Regression to our DBH prediction challenge.
The concept, as Abhishek eventually explained, was beautifully simple. Traditional AI is like a confident student who always gives you definitive answers, even when it’s guessing. Gaussian Process Regression is like an honest student who admits uncertainty - providing both a prediction and a confidence level based on how similar the current problem is to ones it has encountered before.
Here's how we taught our AI to be honest about uncertainty:
Step 1: Collect Some Ground Truth Data: We did send our field team with measuring tapes. But instead of 55,000 trees (which was our original training data), they only visited 7,000 trees. For each tree, we recorded three things: i) Tree Crown size (how wide the leafy part is ii) Tree Height (from drone data) iii) DBH (from wrapping measuring tape around the trunk)
Step 2: The Log Transformation Trick: Here's where it gets mathematically interesting. Trees don't grow in nice, predictable patterns. Some skinny trees are tall, some fat trees are short. The relationship between crown size and trunk diameter is... messy. Abhishek showed us something cool: instead of trying to predict DBH directly, we predict the logarithm of DBH.
"Why?" we asked.
"Because trees follow something called a log-normal distribution," he explained. It's like how most people have average incomes, but a few people are super rich. Most trees have average trunks, but some are really thick. The logarithm transforms this skewed data into a nice, normal bell curve that the math can handle.
Which means, instead of dealing with a skewed distribution, we transform it into a normal distribution that is better suited for the model to learn from.
Step 3: Teaching Uncertainty: This is where Gaussian Process Regression gets magical. Traditional machine learning gives you a single prediction: "This tree's DBH is 15.3 cm."
Gaussian Process Regression gives you a prediction plus uncertainty: "This tree's DBH is 15.3 cm, plus or minus 2.1 cm, and I'm 90% confident it's in that range."
But here's the really clever part - the AI learns when to be uncertain. If you show it a tree that looks similar to ones it's seen before, it'll be confident. If you show it something weird, it'll admit it's guessing.
Abhishek introduced us to two types of uncertainty that sound like Harry Potter spells but are actually crucial for forest monitoring:
Epistemic Uncertainty: Which means "I don't know because I haven't seen enough examples."
Solution: Collect more training data.
Aleatoric Uncertainty: Which means "I don't know because the world is genuinely random."
Solution: Accept that some things are inherently unpredictable.
Our tree data had both types of uncertainty. Some trees were just inherently weird - same crown size, same height, but completely different trunk thickness. Maybe one grew in rocky soil, another in rich earth. Maybe one got more water.
Who knows? The world is random that way.
The Gaussian Process basically learned to say: "For trees with this crown size and height, I've seen trunk diameters ranging from 12-18 cm. Based on the patterns, I think this one is probably 15 cm, but there's natural variation I can't account for."
Here's how we tested whether our AI was actually being honest about uncertainty:
We took our validation data (trees the AI had never seen during training) and asked it to make predictions with confidence intervals. For example:
I'm 60% confident this tree's DBH is between 12-16 cm
I'm 80% confident this tree's DBH is between 11-17 cm
Then we checked: when the AI said it was 60% confident, was it right 60% of the time? When it claimed 80% confidence, was it right 80% of the time?
We plotted these results on what's called a "coverage plot." A perfectly calibrated AI would show a straight diagonal line - 60% confidence = 60% accuracy, 80% confidence = 80% accuracy.
Our AI showed a slope of 1.08 - meaning it was slightly overconfident, but remarkably well-calibrated.
To put it simply, our AI wasn't just making good predictions - it was being honest about when it was guessing.
With honest DBH predictions, we could finally calculate carbon storage for every single tree:
The formula:
Biomass = exp(-1.996 + 2.32 × ln(DBH)) (from Verra’s VMD0001 methodology document)
Total Tree Biomass = Above-ground Biomass × 1.27 [accounting for roots]
Carbon Storage = Total Biomass × 0.5 × (44/12) [molecular weight conversion]
But here's the beautiful part - because our DBH predictions came with uncertainty estimates, our carbon calculations also came with uncertainty. We could say: "This forest has stored 45.3 tons of carbon, plus or minus 3.2 tons, with 90% confidence."
For the first time, we had automated carbon accounting with built-in honesty.
The real test came when we compared our AI predictions to manual field measurements across two case studies (two distinct intervention plots):
Case Study 1: AI predicted 15,480 kg of biomass. Manual measurements (with extrapolation): 18,142 kg. Difference: 15% overestimate.
Case Study 2: AI predicted 5,680 kg of biomass. Manual measurements: 5,693 kg. Difference: 0.2% overestimate.
Our AI was consistently within the uncertainty bounds it claimed. When it said "plus or minus 20%," the real answer was indeed within that range.
And more importantly, we could process entire forests in hours instead of months, and provide uncertainty estimates that manual sampling simply can't match.
By December 2024, we had achieved something quite remarkable: an AI system that could count trees, measure their dimensions, estimate carbon storage and even tell us how confident it was in its own calculations. Fully automated. Completely transparent.
But once we could measure carbon, a different question emerged: what if we could also teach our models to recognize different tree species? Imagine an algorithm that not only says “here’s a tree,” but also tells you this one’s mango, that one’s custard apple, and over there, that’s a bael tree.
The possibilities are powerful. Species detection doesn’t just sharpen carbon estimates - it creates living inventories of trees, mapped by location. Suddenly, it might be possible to get remote data to feed into decisions like setting up a custard apple pulp processing unit in a region with abundant supply, or identifying landscapes where certain species are in decline.
Take the bael tree (Aegle marmelos), revered in Hindu tradition - its leaves offered to Lord Shiva and its fruit turned into delicious candies and jams. Despite its cultural and ecological importance, the bael population in India has shrunk by 25% in just a few decades, landing it on the IUCN’s Near Threatened list.
AI-driven species detection could become a tool not only for planning and economics, but also for conservation, helping us protect what’s sacred before it disappears.
As we began explaining our tech to folks, other on-ground implementers started asking: "Can we use this too?"
We faced a choice. We could keep this as our competitive advantage. Build a business around proprietary forest AI.
Or we could do something different.
We remembered why we started this journey - standing ankle-deep in monsoon mud, watching QR codes wash away, frustrated by how expensive collecting and analyzing high quality and granular data in this sector was. That problem wasn't unique to us. Every nature organization in the world faces similar challenges.
We were also acutely aware of how we'd been able to develop this technology in the first place. It was primarily funded by patient philanthropic capital - funders who believed in us and our vision from year 1 - when the entire F4F team was just the co-founders, one field team member, and a two-person tech team (one of whom was volunteering his time with us). These funders invested not in a guaranteed product, but in a possibility.
We truly think this should be a public good. Given the history, hoarding it as proprietary IP felt wrong. So we have decided to open-source everything.
Named after Anahita, the ancient Persian goddess who was considered a guardian deity, symbolizing life-giving forces, Project Anāhitā is our attempt to democratize forest monitoring. Because, apparently, we needed a suitably grand name for what amounts to years of work that we're now cheerfully handing out to anyone with a conservation and restoration dream.
We're releasing four resources that represent everything we've learned about teaching machines to count trees:
Our complete technical methodology, published as "AI Based Automated Calculation of Carbon Stock in Agroforestry Plantations." Everything from our AI training process to validation results to honest discussions of limitations. The full scientific process behind teaching machines to count trees and estimate carbon storage with uncertainty quantification.
The crown jewel of our data collection efforts: annotated ground truth data with GPS coordinates, real and estimated DBH measurements, species identification, and age estimates for select trees. This dataset represents hundreds of person-hours of our field team measuring trees in 45°C heat. We're sharing it because we believe good training data shouldn't be a competitive moat.
A web interface where any organization can upload drone footage and get automated analysis. No setup required, no technical expertise needed. Just upload your orthomosaic and get back geo-tags, tree counts, carbon estimates, and some species identification.
All our tree detection models, Gaussian Process implementations, and data processing pipelines. It will be available on GitHub with documentation that we hope is better than the usual "it works on my machine" approach.
Ecosystem restoration often fails globally because monitoring is too expensive and accountability is too low. Organizations plant trees, report 'success', and hope for the best because proper monitoring would cost more than the tree planting itself.
We're not trying to build the next unicorn. We're just tired of watching good restoration projects fail because nobody can afford to monitor them properly.
By making professional-grade forest monitoring accessible, we hope to:
Increase survival rates through early intervention when something goes wrong
Improve funding flows through transparent, granular impact data
Scale restoration efforts by reducing monitoring costs from prohibitive to affordable
Build trust in nature-based climate solutions by moving beyond "plant and pray"
Starting with today's research paper, we invite you to dive into our methodology. Poke holes in our approach. Tell us if we got anything wrong.
But we would also love to know if we got something really right!
In the coming weeks, as we release the dataset, platform, and code, we'll walk you through exactly how to use each resource. Whether you're a researcher wanting to build on our models, an implementer needing to monitor your own trees, or just someone curious about how AI can count trees from the sky.
Because the future of forest monitoring isn't going to be built by one organization in central India. It's going to be built by everyone who cares about getting this right.
And maybe, just maybe, we can finally move beyond the "plant and pray" model that's been failing farmers and communities for decades.
In Part #4 of 5 coming in couple of weeks, we'll walk you through our ground truth dataset and share the stories behind some of those measurements. We will also launch TreeLens - our attempt to make drone-based forest monitoring as easy as uploading a photo to Instagram - minus the beauty filters!
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.