RSS Amplifier

softmax · Nov 22, 2024

Humans in the Loop

0
Sign in to vote or save

Mark Redito · softmax

Drawing of a sewing school (1898) by Wally Moes

Imagine an alien intelligence that speaks our language, understands our emotions, and learns from our collective memories. It may seem distant, even artificial, yet at its core, this intelligence is a mirror—a reflection of who we are as humans. Artificial Intelligence (AI) might feel alien, detached, almost beyond our understanding. But in truth, it is anything but. AI is intrinsically human, shaped by our data, our expertise, and our values. The process of creating and interacting with AI is deeply intertwined with human involvement—from the first dataset to the final conversation.

AI is not quite like us. But it has similarities. It knows our language, it has learned how we think, reason, and relate to one another. Across every stage of the AI pipeline, from the inception of data to the responses generated in our chats, humans are in the loop, constantly shaping and guiding the technology. This essay aims to explore the depth of human involvement in AI, showcasing how AI is a continuous loop between human input and machine learning.

It all starts with data. Consider the vast ocean of information that forms the internet: blog posts, social media conversations, tweets, articles, essays, and reviews. Each piece of content—every opinion, every joke, every impassioned argument—is produced by humans. All of it feeds into the AI, becoming the training material from which it learns. This massive dataset, or corpus, is a representation of human thought and language, a mosaic of our collective expression.

Handling this corpus is a task involving many human minds. There are people collecting and curating the data, analyzing it, cleaning it, and ensuring that it is high-quality, balanced, and diverse. The labor of data specialists, engineers, and content reviewers ensures that the model is built on a foundation that is as representative of humanity as possible. Data may be digital, but its origins are undeniably human—a reflection of our voices, our stories, and our beliefs.

Once the data is gathered, the next step is pre-training—a stage that involves the efforts of many researchers and engineers. They are the architects behind the algorithms, meticulously designing the structure of the model. They make key decisions about which architecture to use, which parameters to adjust, and how best to shape the AI's capacity for understanding.

After a training run, these teams evaluate how well the model performs. They analyze metrics, tweak parameters, and iterate again and again until they reach a satisfactory outcome. This process is far from automated—it requires the expertise, intuition, and creativity of people who know how to balance precision with innovation. AI may crunch the numbers, but it is humans who decide what makes a good model—and how best to achieve it.

Fine-tuning is where the human element becomes even more apparent. This stage involves the careful curation of conversations to teach the model how to communicate effectively. Datasets of human dialogues are crafted by experts to show the model what sounds natural, engaging, and coherent. Humans are also responsible for evaluating the model's responses, ensuring that they are not only correct but also relatable and appropriate.

Recently, interesting approaches like using synthetic data—data generated by other AI models—have emerged. However, even these methods require human oversight. People still have to ensure that the synthetic dialogues make sense, have the right tone, and reflect the nuances of human interaction. Beyond technical correctness, there is a responsibility to make sure that the AI aligns with human values—a process that involves creating preference datasets, where evaluators teach the AI what types of responses humans prefer, helping the model be more aligned with human expectations.

Inference is perhaps where AI becomes most tangible to us. This is the stage where most of us interact with models like ChatGPT or Claude—where we type in a question and receive an answer. Here, the model draws on all it has learned from the human-generated corpus and fine-tuned adjustments to provide a response. Millions of people use these chatbots every day, for everything from work to play to companionship. For many, AI has become a daily companion, a helper in navigating the complexities of life.

An interesting phenomenon that occurs during the training of LLMs is the emergence of unexpected capabilities—what researchers call emergent properties. These properties are capabilities the model was not specifically trained to perform, such as solving riddles or demonstrating rudimentary coding skills. While these emergent properties might seem like magic, they are the result of intricate patterns learned from vast amounts of data. This showcases not just the power of the model, but also the limits of human understanding in predicting what an AI might learn. It serves as a reminder that, even with meticulous human oversight, AI can develop capabilities that surprise and challenge us—making the role of humans in interpreting, evaluating, and controlling these behaviors even more vital.

When we talk to an AI, we are closing the loop—we are the final piece of the process that began with data and design. In that moment of interaction, the human experience that was fed into the model now comes back to us, shaped and transformed, yet still undeniably ours. The human touch is there in every line of code, every parameter, every curated response.

From data collection to fine-tuning, AI is an endeavor that is fundamentally human at its core. It learns from our words, adapts based on our preferences, and becomes a companion—reflecting the best and worst of our human condition. It is a loop that begins and ends with us, shaped by our creativity, our labor, and our desire to build something that extends beyond our natural capacities. To think of AI as detached from us would be to ignore the countless human minds and hearts behind every interaction. At every step, humans are in the loop, ensuring that this alien intelligence remains, at its heart, a reflection of who we are.

While AI development is deeply rooted in human effort and creativity, we must also acknowledge some emerging challenges that come with the increasing prevalence of AI-generated content. As AI becomes more capable, the lines between genuine human data and synthetic data blur, leading to what some call "AI Slop"—a scenario where the abundance of AI-generated content starts to dilute the quality and authenticity of the data used for training. This growing mixture of synthetic and genuine human experiences may impact the richness of future AI models, potentially creating feedback loops where models train on lower-quality, AI-generated material rather than original human content.

Moreover, the automation of training and evaluation processes raises concerns about losing the nuanced human judgment that is crucial to maintaining quality and ethical standards. As we move forward, it is important to remain vigilant in preserving the value of human input—ensuring that AI continues to reflect true human diversity, creativity, and depth. By recognizing these challenges, we can make informed choices that keep humans firmly in the loop, guiding AI toward a future that remains connected to our best qualities.

🌀

Share

Read the original on redito.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.