RSS Amplifier

Fellows Fund · May 30, 2024

Gen AI Summit 2024 | Gen AI Startup Panel

0
Sign in to vote or save

Clarice Wang, Eric Yu · Fellows Fund

  • Jia Li, Co-Founder, Chief AI Officer & President, LiveX AI Inc.

  • Dmytro (Dima) Dzhulgakov, Co-Founder & CTO, Firework AI

  • Hassaan Raza, Co-Founder & CEO, Tavus

  • Young Zhao, Co-founder & CEO, Opus Clip

  • Charles Elkan, Venture Partner at Fellows Fund; Ex-Global Head of ML, Goldman Sachs; Professor of CS, UCSD (moderator)

The Gen AI Startup Panel, moderated by Charles Elkan of Fellows Fund, featured Jia Li (LiveX AI), Dmytro Dzhulgakov (Firework AI), Young Zhao (OpusClip), and Hassaan Raza (Tavus). OpusClip converts long-form videos into engaging short clips for social media, LiveX AI develops multimodal AI agents, Firework AI customizes and deploys open-source generative models, and Tavus creates digital video twins and avatars. They emphasized the transformative impact of AI and highlighted the challenge of aligning AI product capabilities with customer needs. Key themes included the importance of listening to customers while managing expectations amid AI hype. The panel also addressed ethical considerations, technical challenges like cost and scalability, and predicted a shift towards multimodal models for comprehensive understanding.

Introduction: Our next Fireside Chat focuses on Gen AI startups, featuring Jia Li, Co-Founder, Chief AI Officer and President of LiveX AI; Dmytro Dzhulgakov, Co-Founder and CTO of Firework AI; Young Zhao, Co-founder and CEO of Opus Clip; and Hassaan Raza, Co-Founder and CEO of Tavus. This session is moderated by Charles Elkan, partner at Fellows Fund and founder of Ficc.ai. Let's give them a warm round of applause as they come up on stage.

Charles Elkan: Thank you all for being here. It's a pleasure to host this panel and get to know these impressive founders building amazing products. I have some questions prepared, and I'm looking forward to hearing their answers. First, I'll ask the panelists to briefly explain how they use AI in their companies and share some insights into their entrepreneurial journey.

Young Zhao: Hi, everyone. I'm the co-founder and CEO of OpusClip. We turn long-form video footage into social, engaging short clips with just a click. We leverage various video understanding and generation technologies to help anyone tell a great story on their social media. I'm happy to share more in future discussions.

Jia Li: Hi, everyone. I'm Jia, an AI technologist. My company, LiveX AI, builds multimodality AI agents for everyday living. Gen AI is transforming how we interact with products, surveys, hardware, and devices. Currently, we are building a business AI agent customized for businesses to serve their customers. In the future, we hope to extend this to AI agents that can serve everyone personally in their everyday life.

Dima Dzhulgakov: Hi, I am the CTO of Firework. Firework is a developer platform that helps developers and businesses build with open-source generative models across all modalities. You’ve probably heard about models like Llama, Stable Diffusion, and Whisper. We help people customize those models for different use cases. Whether you have customized models or just the base model, and want to deploy a new application, we take care of all the serving, scaling, and performance needs, allowing you to build your application quickly and at a low cost. We pride ourselves on being one of the fastest providers on various leaderboards. We focus on customization for different use cases, both in terms of quality through fine-tuning and performance by tailoring deployment configurations. So if you have low latency needs for chatbot systems, that's our focus.

Hassaan Raza: Hey, everyone. I'm the founder and CEO of Tavus. We're a generative video research company that builds models for creating personal digital twins. We provide these models to developers and enterprises to offer avatar and cloning features within their platforms. This can help you create realistic videos from just text without recording, or help companies create videos in languages that a person doesn't speak. For example, healthcare software companies use it to help doctors deliver video messages in the patient's native language. We even have models for real-time use cases, like sending your avatar into a Zoom call for you. We focus on building high-fidelity, easy-to-use cloning models ethically and thoughtfully, with consent only. Ultimately, we’re helping customers leverage their digital likeness at scale, enabling them to reach out to far more customers and users than ever before.

Charles Elkan:Very impressive, thank you. Let's start with some higher-level questions about being a startup founder. As a founder, you need to continuously sell your vision—to investors, employees, and customers. Which of these is the most important and which is the most difficult? More generally, what have you learned about being successful in sales? Let's start with you, Hassaan.

Hassaan Raza: Good question. I'd say the most challenging is selling your vision to customers. With investors, you're selling them a dream, your dream. But with customers, the dream has to be rooted in the realities of their business and what your product is today, not its future state. Selling to customers requires tailoring your vision to their specific needs, which can be challenging. Often, you're thinking of a future state with many features and functionalities, but you have to deliver what you can today. What has been successful for us is partnering with customers and making them feel as if their roadmap is our roadmap. Aligning with them and helping them build brings our dream into the context of their business. This customer-first approach has been key in closing deals, by being so invested in their roadmap and success.

Charles Elkan: Jia, it looks like you have a perspective on this.

Jia Li: Our customers are typically less technical, unlike many investors who have caught up with or invested a lot of time into understanding Gen AI. This makes it challenging to convey the value that Gen AI can bring them today and its bigger potential down the road. We've learned that in the selling process, it's crucial to listen to customers' pain points and understand their challenges. It aligns much better if you can help them solve their problems.

Charles Elkan: Young, Dima, do you have a perspective you'd like to add?

Young Zhao: Telling the mission to employees, customers, and investors is quite different in our case. It's pretty easy to convey our vision to customers because we are a job-to-be-done-focused company with a product that is easy to understand and solves customer pain points. Many customers can articulate our vision better than we can, especially influencers on social media. However, the most important is conveying the vision to employees. We are building a product with no benchmark, solving problems with purely innovative solutions. It's critical to keep iterating, re-emphasizing, and communicating the vision to ensure everyone is on the same page. Telling the vision to investors is challenging because many investors are not creators and have a hard time understanding our space and use cases. Some investors are really into social media, but the vast majority are not, making this the most challenging aspect for us.

Dima Dzhulgakov:  For a developer-focused product, I definitely agree that focusing on customers is the most important. Building something people want makes it easier to motivate employees and show investors the value. There are two components to selling a developer product: selling to the people who will use it hands-on, where small details matter a lot, and addressing the needs of the customer as a business. Understanding their end use case and helping them reframe problems can add significant value. It's about addressing their business problem, with the technical piece being just a component of that. Going the extra mile to help customers along the way creates a lot of value.

Charles Elkan: On the theme of customers and product, how do you think about customers versus competitors? Do you want to be unique? Or do you want to be addressing a huge market? Or do you have a different way of thinking about the question?

Dima Dzhulgakov: I would say the market is big, and the general space of applications for Gen AI is much, much bigger than even the previous wave of deep learning. I was a core developer on PyTorch for many years, and there was a huge market for developing deep learning applications, but it required much more technical talent. The applications for foundational models in this next wave represent an even larger market. It's much more important to focus on customers and solving their problems because the market is big enough that you shouldn't worry about competitors at this point. The best product ultimately wins, so if you build the right product for the right customer and ensure proper distribution, competitors shouldn't be a major concern.

Jia Li: Right now, there's a lot of talk and fancy demos around Gen AI for customer experience and AI agents. This can create misunderstandings from the customer's perspective, making them think they can achieve something unrealistic today. Building a demo might take 5% of the time, but running the last mile to build a full product takes 95% of the effort. It's crucial to cut through the noise, educate customers, and build a trustworthy relationship with them.

Hassaan Raza: For us, we've transitioned from building models for our own sales and marketing product to focusing on letting these models be used by developers. The generative AI space is novel and new, and we see a large advantage in powering various products and ideas by providing the models rather than keeping them for our own product. This way, we can leverage the creativity and innovation of developers across different fields.

Young Zhao: In our use case, we started with a niche, being the best in the breed early on. Once we were confident in our space, we began gradually expanding our use cases beyond our initial customer base and long-to-short editing. Ultimately, we believe the market for social video creation and editing is huge. For startups, it's crucial to be the best product even as you expand your use cases and technologies. You can't go too crazy, especially with a limited team. We are working on the next generation of our product, pacing our expansion carefully—three times, five times, ten times—into a bigger market.

Charles Elkan: Thank you. So changing gear a little bit. Gen AI. Obviously, it's a platitude today, it's a very powerful technology, and we need to be sure that we use it for good. So the question I'd like to ask is, what's an example of a real-world ethical issue that you've encountered at your company? And how have you addressed it? Hassaan, I know I asked you this question when we were talking just yesterday.

Hassaan Raza: For us, there's always the question around disclosure and consent. We're building models to create digital twins, which means we're cloning your likeness—your face and your voice. We've always taken the approach that consent is the bare minimum. You have to get consent; we will not build models that allow you to create a clone of someone else. This does come with a trade-off. For example, there's the opportunity for virality—you could post a video of someone very famous, create a clone of them, and generate a lot of great marketing hype. But we're always saying, hey, it's just not worth it. We're going to stay on the side of consent. We're not going to allow anyone, or ourselves, to clone someone without their consent. So that's a trade-off that we face internally.

Dima Dzhulgakov: From an infrastructure platform side, I definitely see a lot of customers trying to figure out the right balance. There's a big gap between just getting the model trained and actually turning it into a product. That's where a lot of responsible and ethical issues come in. Technically, it means that guardrails, safety filters, and all that stuff also need to be built in. Sometimes people see model training as the consuming part—it might be the most resource-intensive part, consuming GPUs—but actually, in the employee development phase, the safety and responsibility part takes quite a big piece. As a platform provider, we take the privacy of data very seriously. We don't log data unless it's necessary for the customer's specific need of improving and fine-tuning their model. Logging data and retention, and being responsible on that side, is our contribution to this. But I think the main ethical issues are closer to the product side, and how to enable responsible use is something we're focused on.

Young Zhao: In our use case, we've encountered millions of copyright issues. A lot of our users were using other people's content, especially in the early days when we didn't do anything to prevent it. We heard a lot of complaints from big influencers and creators, asking why they could see their content everywhere on YouTube with our captions and animations. We've been very proactive in addressing this. First, we've been warning our users that if they use others' content, their accounts will be banned. Over time, we've seen fewer people using others' content because platforms like YouTube and TikTok also address this issue. Additionally, we're constantly revamping the user experience workflow to make it harder for users to infringe on copyright. We tailor the workflow for legitimate users and create friction for those we want to prevent. That's how we deal with this issue.

Jia Li: For us, the public perception of job replacement is a significant ethical issue. When an AI agent becomes more efficient and advanced, the media often frames it as job replacement. In our experience, when we interact with customer support teams or pre-sales teams, many initially feel guarded, thinking AI will replace them. But in reality, many repetitive jobs are not loved by anyone—long hours and repetitive answers. Many are much happier when AI can help them solve challenging problems, allowing them to handle more complex situations better. I believe it's going to be about job transformation instead of job replacement. With Gen AI being such a powerful tool, everyone will learn and adapt to it.

Charles Elkan: Thank you. That's great to hear. So now, I know we have many technical people in the audience and very technical founders on stage here. Let's change gears and ask a couple of more technical questions. You're all bringing AI to production, and we know there are many challenges in that. What are the toughest challenges you've found in bringing AI to production and scaling it up when you get traction with users? I'm thinking of cost, latency, reliability, hallucinations, prompt engineering, guardrails, or something else.

Dima Dzhulgakov: Short answer: all of the above. The journey of production in machine learning or Gen AI is a progression. Initially, it's about getting something to work and iterating quickly. If it's prompt engineering, iteration is easy. If it's fine-tuning, then job training must be super quick. The cost in the initial phases isn't a big issue since you're running on a small scale to build a demo or prototype to validate the product. As you scale up, cost and performance, like latency, become critical. That's when specialization or smaller models through fine-tuning helps. Customizing deployment options to meet specific use cases, such as tight latency requirements for financial systems, is crucial. Infrastructure optimizations tailored to use cases can significantly improve efficiency without changing the model. The challenges shift from iteration speed to quality, scaling, and application integrity.

Jia Li: Building a multimodal AI agent, latency is the most challenging aspect. Most AI agents today are used for generating reports or surveys where users can wait, but we need real-time AI agents with multimodal input and high accuracy for mission-critical tasks. Collaborating with Nvidia, we've optimized hardware and software to speed up token processing over six times. Still, more advancements in hardware, memory, and edge deployment are needed for real-time, personalized AI agents.

Hassaan Raza: I agree. Different stages of the development cycle bring different challenges. Initially, during the research phase, controlling hallucinations, implementing safeguards, and ensuring stability are crucial as models aren't always exposed to real-world scenarios. During production, scale and latency become more critical, along with cost efficiency. Extensive testing helps handle edge cases that aren't apparent in research environments. Each stage has its unique focus, but all these aspects are important throughout the process.

Charles Elkan: Let me move on to a related question. Explainability in AI models is a hot topic. How important have you found explainability to be for your users and customers, and how do you provide explanations?

Young Zhao: Our product is straightforward as we generate videos that are visual and understandable. We emphasize the explanatory aspect by allowing users to preview and replay results to ensure they want the video. We also have a feature called "virality score" that predicts how viral a video might go on social media. This model constantly tracks performance and helps users understand the potential of their content. Our users love this feature, even more than detailed explanations of the AI's workings.

Hassaan Raza: Explainability is crucial, especially for customers using our product within their own platforms. Predictability in model performance and understanding why things go wrong are vital. Clear messaging around training failures or unexpected outputs is essential. Providing insights into proper prompting and ensuring good results is key, especially for enterprise customers.

Dima Dzhulgakov: Explainability varies by product importance. For LLMs, you can interact in natural language, and techniques like chain-of-thought reasoning help by breaking down problems into pieces, improving quality and offering insight into decision-making structures. From an infrastructure perspective, reproducibility is vital. These models are stochastic, and getting consistent results across different hardware is challenging. Ensuring reproducibility for debugging downstream workflows is essential.

Jia Li: Our customers care more about the process than the model's inner workings. They want insights from AI interactions, like popular topics or top questions, to inform product design or upgrades. Gen AI is perfect for analyzing data trends and delivering actionable insights, which is what our customers value.

Charles Elkan: Moving on, you're all building on foundation models. What do you see as the future of foundation models? Will they all be multimodal and larger, or is there a different direction you'd like to see?

Hassaan Raza: The future of foundational models is multimodal. With the recent advancements, we see the importance of understanding different modalities for more powerful and insightful models. The core foundational models in the coming years will naturally be multimodal, reducing the need for translation features, which often lose context.

Dima Dzhulgakov: Absolutely. Multimodality is coming, and we'll see advancements in spatial reasoning and video generation. For embodied AI and robotics, integrating different sensors will be crucial. There will always be top-tier, jack-of-all-trades models, but many applications won't need all modalities due to cost or latency constraints. Customizable, smaller models for specific use cases will also be important.

Young Zhao: There’s consensus that models will become more multimodal. However, the current models still struggle to fully understand images or videos, akin to reading with imagination alone. Future foundational models need better ways to understand content accurately. They should also better support specific applications and use cases, enabling developers to leverage foundation models for targeted fields more effectively.

Jia Li: I strongly believe in the multimodal agent aspect. There's still a significant gap in multimodal capabilities because the data used to train these models is limited by past technology paradigms. There's valuable human behavior and interaction data that hasn't been tapped yet. Future models and hardware form factors will benefit greatly from this untapped data.

Charles Elkan: This has been a fascinating discussion. Unfortunately, we have less than a minute left. I'll squeeze in one more question. Are there limits to the intelligence of transformers, or are we not close to that ceiling yet? Is it the quality and quantity of data that's the limiting factor?

Dima Dzhulgakov: Transformers have won the hardware lottery and have been optimized extensively. They’re scalable, and we’ll see a few more orders of magnitude in their capabilities. However, incorporating external information and ensuring reliable reasoning remain challenging. While transformers will continue to scale, they might not achieve AGI in the broadest sense without significant breakthroughs.

Hassaan Raza: I agree. There's still room for growth in transformers, particularly in data collection and optimization. Replicating human thought processes and workflows will challenge transformers' limits. Although they might not achieve AGI soon, there's much room for improvement within the current architecture.

Jia Li: We've focused a lot on models, but data quality and quantity are equally important. The data used to train these models often comes from internet sources, which isn't always optimal. High-quality, abundant data is crucial, and future model architectures will likely evolve with better data.

Young Zhao: I’m not an expert on this, but I agree with the comments. Current models can't fully replicate human perception. There's significant room for improvement, particularly in addressing more complex tasks. Future advancements should focus on making models capable of higher-level tasks, moving from junior-level to senior-level performance, bringing us closer to AGI.

No posts

Read the original on fellowsfund.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.