RSS Amplifier

AI Cathedral · Nov 7, 2024

Leading AI Engineering Teams

0
Sign in to vote or save

Nik Pash · AI Cathedral

When I talk to engineering leaders struggling with their AI teams, I often hear the same complaint: Why is progress so slow? We’re pouring resources into models, pipelines and trendy libraries, but it feels like we’re shooting in the dark—there’s no clear path to a 'finished' product.”

This dissatisfaction comes from a fundamental misunderstanding: AI development is not just engineering; it’s applied research. This fundamentally alters our understanding of progress, goals, and team management. In my previous article, I talked about using clear, metrics-driven communication to drive improvement in AI systems. Today, I want to focus on how to effectively lead an AI team.

The ticket isn’t the feature; it’s the experiment. The real deliverable? Learning.

Note: Who am I? I'm an AI engineering consultant and founder specializing in process improvement for rapidly growing startups. With a ML background on Meta’s Knowledge Graph team, I've spent the past year collaborating with numerous AI startups. This experience has given me unique insights into common pitfalls and best practices in the field. My goal is to help engineering teams avoid recurring mistakes and implement more effective, data-driven processes for AI development and deployment.

Traditional software development is deterministic: if you want to build a login feature, you know exactly what success looks like. The path might be complex, but it's clearly defined.

AI development exists in a different paradigm. When we want to "improve accuracy" or "reduce hallucinations," we aren't implementing a known answer; we're looking into whether a solution even exists. It is more similar to scientific inquiry than traditional engineering:

Engineering Ticket:

"Implement a magic link login feature" - Well-defined specifications, established best practices, and a definite result - Success is binary, and the timeline is known.

AI Research Ticket:

"We need to improve the quality of our conversations between assistants and users" - Uncertain if completely solvable; several conflicting strategies; probabilistic results; a spectrum of success; an unclear timeline

This fundamental difference means we need to radically rethink how we measure progress and set goals, while also reconciling this with conventional product needs. Here's why:

  1. Progress isn't linear – In conventional engineering, halfway through a timetable usually indicates halfway done. In research, you could go weeks without making any progress, just to have a breakthrough. Alternatively, demonstrate that your entire method will fail.

  2. Success Isn't Guaranteed – Sometimes the research simply does not support our product aims. No amount of engineering brilliance can alter the fundamental mathematics or constraints of our current methods.

  3. Speed of Learning > Speed of Development – The most important metric is not how rapidly we can implement features, but how quickly we can test hypotheses and learn from the outcomes.

This basic distinction means we need to dramatically rethink how we set goals and measure progress.

With all this in mind, here’s what this entails for your AI team:

Rather than asking, "Why isn’t conversation quality improving?" ask:

  • How many hypotheses have we tested this week?

  • What's our experiment iteration speed?

  • What are we spending time on between experiments?

  • What's stopping us from running more experiments?

Each experiment, successful or not, should improve the team's intuition about:

  • What approaches are likely to work

  • What the fundamental limitations might be

  • Where the unexplored territories are

  • Which paths are dead ends

Leaders need to get comfortable with updates like:

"We conducted five experiments this week. All of them failed, but we eliminated three key approaches and found a potential new direction. "Our hypothesis space is narrowing."

This is what progress in AI research looks like.

In research-heavy AI projects, it’s easy for teams to feel disconnected from a shared purpose, especially when experiments bring ambiguous results. This is where having a North Star metric can make a massive difference. A single, overarching metric connects every experiment, failure, and breakthrough to a tangible outcome that everyone on the team understands and strives toward.

  1. Identify a Metric That Matters to Users

    Choose a metric that represents the value your AI brings to the end-user. For example, if your goal is to improve conversation quality, you might choose a metric like “percentage of conversations with positive feedback.” This metric reflects value in terms of the user’s experience, not just internal measurements like model accuracy.

  2. Break Down Key Drivers

    Once you’ve established your North Star, identify which inputs or processes the team can influence to move that metric. For a conversational AI team, these could be the number of contextually accurate responses or the reduction in misunderstood queries.

  3. Integrate the Metric into Team Routines

    Make your North Star part of every meeting, weekly review, and sprint planning session. When everyone knows what they’re aiming for and how each experiment feeds into that metric, it transforms abstract research goals into actionable steps. This alignment allows the team to connect their daily work to real business impact.

  4. Create a Feedback Loop Around the Metric

    Implement regular reviews where the team can analyze how recent experiments impacted the North Star metric. Even if some experiments fail, every test brings new data that drives collective learning and hones intuition about what works.

By aligning around a North Star metric, your team gets a powerful, unifying sense of purpose. This metric helps unify your team and give meaning to their experiments, serving as a rallying point even in the ambiguity of research.

When you embrace the research mindset, different questions emerge:

Instead of:

"When will the conversation quality improve?"

Ask:

"What's preventing us from running more experiments to improve conversation quality this quarter?"

Instead of:

"Why isn't this working yet?"

Ask:

"What have we learned about what doesn’t work and why?"

Instead of:

"Can we speed this up?"

Ask:

"What's preventing us from experimenting faster?"

Here are the metrics I monitor that help me make decisions:

  • Number of hypotheses examined each week

  • Average time per experiment

  • Quality of documentation from failed experiments

  • Growth in team's understanding of the problem space

  • Time spent waiting for compute/data

  • Reproducibility of experiments

  • Modularity of systems to easily swap in/out things during testing

  • Time to test new hypotheses

Instead of:

"Improve accuracy by 10% this quarter"

Try:

"Run 15 experiments per month, with each experiment having clear success/abandon criteria and documented learnings"

This changes the conversation from "Why aren't we succeeding?" to "How can we run more experiments? What's slowing us down? How can we learn faster?"

The most difficult aspect of AI development is balancing research reality and product requirements. Here's how I tackle it:

  • Research metrics: experiment velocity and learning rate.

  • Product metrics include accuracy, delay, and cost.

  • These metrics work together: while research metrics drive learning, the North Star ensures alignment with product impact, bridging the gap between experimentation and end-user value.

  • Set clear decision points

  • Define minimum viable improvements toward the North Star

  • Have backup plans for insufficient progress

  • Keep exploring alternative approaches

  • Build modular systems that can swap approaches

  • Document learnings for future attempts and keep the team focused on how each step could ultimately impact the North Star

By grounding both research and product goals in a shared North Star metric, you ensure that even experimental work remains connected to the team’s overarching objective, helping bridge the gap between research uncertainty and product impact.

Leading AI teams effectively requires a shift in mindset from traditional engineering to applied research. Here are the key takeaways:

  • The Ticket is the Experiment: In AI, tasks aren’t features but experiments, with learning as the main outcome.

  • Embrace Non-Determinism: AI work is probabilistic, not binary; success isn’t guaranteed, and progress is rarely linear.

  • North Star Metric: Define a single, user-centered metric to unify the team’s work and link experiments to real value.

  • Focus on Experiment Velocity: Track the number of hypotheses tested, experiment iteration speed, and time between experiments.

  • Prioritize Learning Over Development Speed: Your team’s understanding of what works (and doesn’t) matters more than just shipping features.

  • Balance Product and Research Goals: Use dual metrics—like accuracy for product goals and experiment velocity for research—to keep AI development focused on both immediate value and long-term learning.

  • Foster Autonomy Through Clarity: With a clear North Star and regular feedback, teams make better decisions independently, allowing for adaptive problem-solving and innovation.

With these principles, AI teams can work purposefully, grounded in learning and aligned with real business impact.

PS: What do you think is your team's largest impediment to functioning this way? Leave a comment below.

Leave a comment

No posts

Read the original on pashpashpash.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.