RSSAmplifier

Blog

A Computer Scientist in a Business School

Random thoughts of a computer scientist who is working behind the enemy lines; and lately turned into a double agent.

behind-the-enemy-lines.comRSS feed ↗25 posts

Latest posts

The Lawyer I Never Hired: How ChatGPT Got Me $4,760 Out of Norse Atlantic

TL;DR: Our Norse Atlantic flight out of Oslo was delayed overnight for technical reasons. Norse offered a $25 refreshment card and then ignored our claims. Because the flight departed the EEA, EU261 applied even though we live in New York. I had no idea how to enforce a European regulation against a Norwegian company from 6,000 kilometers away. ChatGPT-Pro did. It walked me through the national…

Taste Is Not Enough. Reality or Bust

A paper published in Nature on March 25, 2026 describes "The AI Scientist," a system built by Sakana AI that automates the full cycle of scientific research: idea generation, experiments, analysis, writeup, even peer review submission. The marginal cost of producing paper-shaped research output is collapsing. So, when paper production becomes cheap, what is next? 100,000 axiom systems and counting…

How I Stopped Being a Copy-Paster for My AI Agent: Claude Code, Google Cloud, and the Loop to Close

TL;DR: Your AI agent in Claude Code on the Web can use Google Cloud (or AWS/Azure) to store large datasets, run long computations, deploy web apps, and schedule recurring jobs. Once you have a cloud account and project, the repo-specific setup takes about five minutes: Set an encryption password in your environment settings (see Step 1 below). If you only use one cloud provider, name it…

"Let's Work on the Next Task": Claude Code, GitHub, and the Most Diligent Project Manager I've Ever Had

In my previous post , I described how working with AI agents felt like managing an infinitely large, infinitely diligent team. I wrote about pairing Claude with GitHub, giving it context files and task lists, and watching it come back with actual deliverables. After that post, I got questions from a lot of people asking how to actually set this up. Even from people I assumed were already using…

Listening to My Students at Scale: Exit Tickets, NotebookLM, and the Tightest Feedback Loop I've Ever Built

It started at a teaching workshop, last semester: Craig Kapp and Rob Egan presented a seminar at the NYU Center for Teaching and Learning called "Real-Time Insights: Leveraging AI for Responsive Teaching in Large Classrooms." They (re-)introduced a deceptively simple concept: the exit ticket . The idea is that at the end of every class session, you ask students three quick questions, each with a…

Everybody Is a CEO Now (And What Exactly Am I Doing Here?)

It's hard to pinpoint the exact moment when something fundamentally shifts. There's no day when you wake up and say, "Today, everything is different." It's more like boiling a frog. Except in this case, the frog is me, and the water feels amazing . Over the last few weeks, a confluence of AI developments crossed an invisible threshold. None of them is dramatic on their own. All of them, together,…

Fighting Fire with Fire: Scalable Personalized Oral Exams with an ElevenLabs Voice AI Agent

Paper: This blog post has been expanded into a full paper: Scalable and Personalized Oral Assessments Using Voice AI (Panos Ipeirotis and Konstantinos Rizakos, Communications of the ACM, forthcoming ). The paper includes the complete system design, failure mode analysis, student experience data, and all prompts as appendices. Demo Video: A YouTube video showing an example of how a student would…

Training LLaMA using LibGen: Hack, a Theft, or Just Fair Use?

Imagine you're building a Large Language Model. You need data—lots of it. If you can find text data of high quality, vetted, truthful, and useful, it would be... great! So, naturally, you head online and find a treasure trove of books neatly indexed, conveniently downloadable, and completely free. The catch? You're looking at LibGen —one of the most infamous pirate libraries on the internet. This…

Copyright, Fair Use, and AI Training

[We tested the o1-pro model to give us a detailed analysis of the legal landscape around copyright and the use of copyrighted materials to train LLMs. The full discussion is available here. Below you will find a quick attempt to summarize the (much) longer report by o1-pro.] What is Copyright? (And Why Should You Care?) Imagine you spend months writing a book, composing a song, or designing a…

Developing Grading Rubrics using Docent

When I explain the concept of Docent, a common first question I hear is if AI grades assignments, can't students just use AI to do their homework? They imagine a scenario where professors create assignments with AI, students complete them with AI, and graders assess them with AI as well. But, let me clear that up—Docent doesn't work like that. While we were crafting Docent, we figured out we…

Grading with AI: Introducing Docent

TL;DR An alpha version of Docent, our experimental AI-powered grading system, is now available at https://get-docent.com/ . If you're interested in using the system, please contact us for support. The Challenge of Grading One thing that I find challenging when teaching is grading, especially in large classes with numerous assignments. The task is typically delegated to teaching assistants with…

The PiP-AUC score for research productivity: A somewhat new metric for paper citations and number of papers

Many years back, we conducted some analysis on how the number of citations for a paper evolves over time. We noticed that while the raw number of citations tends to be a bit difficult to estimate, if we calculate the percentile of citations for each paper, based on the year of publication, we get a number that stabilizes very quickly , even within 3 years of publication. That means we can estimate…

Tell these fucking colonels to get this fucking economist out of jail.

Today is October 18th. It is 41 years since Greece voted for Andreas Papandreou with a 48% vote percentage to be elected as prime minister, fundamentally changing the course of history for Greece. Positively or negatively, this is still debated, but the change was real. On October 6th, Roy Radner passed away at the age of 95. He was a faculty member at our department and a famous microeconomist…

"Geographic Footprint of an Agent" or one of my favorite data science interview questions

Last week we wrote in the Compass blog how we estimate the geographic footprint of an agent . At the very core, the technique is simple: Use the addresses of the houses that an agent has bought or sold in the past; get their longitude and latitude; and then apply a 2-dimensional kernel density estimation to find what are the areas where the agent is likely to be active. Doing the kernel density…

Mechanical Turk, 97 cents per hour, and common reporting biases

The New York Times has an article about Mechanical Turk in today's print edition: " I Found Work on an Amazon Website. I Made 97 Cents an Hour ". (You will find a couple of quotes from yours truly). The content of the article follows the current zeitgeist: Tech companies exploiting gig workers. While it is hard to argue that there are tasks on MTurk that are really bad, I think that the article…

Distribution of paper citations over time

A few weeks ago we had a discussion about citations, and how we can compare the citation impact of papers that were published in different years. Obviously, older papers have an advantage as they have more time to accumulate citations. To compare papers, just for fun, we ended up opening the profile page of each paper in Google Scholar, and we analyzed the paper citations years by year to find the…

How many Mechanical Turk workers are there?

TL;DR: There are about 100K-200K unique workers on Amazon Mechanical Turk. On average, there are 2K-5K workers active on Amazon at any given time, which is equivalent to having 10K-25K full-time employees. On average, 50% of the worker population changes within 12-18 months. Workers exhibit widely different patterns of activity, with most workers being active only occasionally, and few workers…

Why was my Amazon Mechanical Turk registration denied?

( This is my answer to a question posted on Quora ) Mechanical Turk is a platform for work. Workers get paid, which makes now Amazon a payment processor. Payment processors are moving money on behalf of other people, and therefore are under heavy scrutiny from the US government for issues related to money laundering (AML), counter-terrorism, tax compliance, etc. One of the key things that is…

AlphaGo, Beat the Machine, and the Unknown Unknowns

In Game 4, of the 5-game series between AlphaGo and Lee Sedol, the human Go champion, Lee Sedol managed to get his first win . According to the NY Times article: Lee had said earlier in the series, which began last week, that he was unable to beat AlphaGo because he could not find any weaknesses in the software's strategy. But after Sunday's match, the 33-year-old South Korean Go grandmaster, who…

A Cohort Analysis of Mechanical Turk Requesters

In my last post, I examined the number of "active requesters" on Mechanical Turk, and concluded that there is a significant decline in the numbers over the last year. The definition of "active requester" was: " A requester is active at time X if they have a HIT running at time X ". A potential issue with this definition is that an improvement in the speed of HIT completion (e.g., due to increased…

The Decline of Amazon Mechanical Turk

It seems that after years of neglect, Mechanical Turk is starting to lose its appeal. In our latest measurement, we see Mechanical Turk losing 50% of its requesters in a YoY measurement. A few days ago, Kristy Milland (aka SpamGirl) asked me if there is a way to see the active requesters on Mechanical Turk over time. I did not have this dashboard on Mechanical Turk tracker, but it was an important…

An API for MTurk Demographics

A few months back, I launched demographics.mturk-tracker.com , a tool that runs continuous surveys of the Mechanical Turk worker population and displays live statistics about gender, age, income, country of origin, etc. Of course, there are many other reports and analyses that can be presented using the data. In order to make it easier for other people to use and analyze the data, we now offer a…

Postdoc Position for Quality Control in Crowdsourcing

The Center for Data Science at NYU invites applications for a post-doctoral fellowship in statistical methodology relating to evaluating rater quality for a new research program in the application of crowdsourcing ratings of human speech production. Duties and Responsibilities : This is a two-year postdoctoral position affiliated with the NYU Center for Data Science. The successful candidate will…

The World Bank Report on Online Labor

I am often asked about statistics and data about the global population of "crowdsourcing" workers, going beyond Mechanical Turk. I am happy to say that from now on I will be able to point everyone to a study from The World Bank , which I was fortunate to participate in. The report examines the global landscape of online labor, identifying the opportunities and providing statistics about the global…

Demographics of Mechanical Turk: Now Live! (April 2015 edition)

One of the most common questions that I receive is whether I have new data about the demographics of Mechanical Turk workers. The latest data that I had collected were back in 2010, and it was not clear how things have changed since then. The key problem was not that I could not run additional surveys; that would have been trivial. However, the results of the surveys were always changing over…