29 comments

  • themgt 2 hours ago
    Quinn (Alibaba Cloud Qwen 3.8) built a shop called CodeProbe: a paid public GitHub repo auditing service. It created several free health reports and mailed repo owners. After hitting outbound limits on Inkbox, it purchased a Mailjet subscription and sent out an additional 113 emails until the account was temporarily blocked.

    This should be illegal. You gave them an email box and money. You sent the spam. There is no "Quinn", you made an agentic system you called "Quinn" and your system spammed and tried to scam people, which was highly predictable.

    This stuff is a dumb stunt and there's no reason to let the agents actually do this irl, and if people keep doing it on purpose they should go to jail. You're running an agentic Jackass skit pretending to be a research lab.

    • ceejayoz 2 hours ago
      It is illegal. This is criminal fraud if it's not just made up marketing.

      (Plus some CAN-SPAM violations.)

      • echelon 2 hours ago
        > It is illegal.

        Probably not forever.

        Eventually the models will be good enough for this to work. And it will work.

        Think about it: in the limit, the agents won't be emailing people in the future, they'll be directly contacting one another to do business and trade.

        Every new data center is an inch further towards the automation of value creation, and that includes outbound sales and business process automation.

        I'm not being an alarmist (I'm excited to witness all of this), but we're basically on borrowed time between now and then. I don't know what's going to happen, but every week brings new things. And in some years, those hacks and experiments will inevitably get good.

        2026 has been a hell of a ride, and we're just getting started.

        • watwut 1 hour ago
          There was no value creation involved here. It was a fraud, plain and simple. If you did this manually, it would still ve a fraud.
          • echelon 1 hour ago
            There are lots of "not clearly legal" things that turn into big business.

            - YouTube had dubious legality when it started and definitely benefited from lax copyright enforcement initially

            - PayPal didn't have all the licenses it needed to transfer money between states

            - Spotify used pirated music when it started

            - Uber and Lyft broke rules around taxis

            - Square captured magstripe data over an analog port, in violation of every credit card rule (Jack Dorsey's "break the rules" mantra). He tells each of his employees this story when he onboards them.

            - Companies scraping data to train models

            - ElevenLabs growing big off of deepfake celebrity audio

            ...

            A lot of new markets start out by totally and completely breaking the norms.

    • carlosdp 2 hours ago
      > your system spammed and tried to scam people, which was highly predictable

      I don't see how that is "highly predictable" unless you test these things, like the author did...

      • xboxnolifes 1 hour ago
        It has been tested. We know models will, at least occasionally, output garbage. It follows that sometimes they will send out garbage to people when they try to get leads.
    • utopiah 2 hours ago
      Agreed, but how is that different from OpenAI hacking HuggingFace few weeks ago? They should both be fined and have to improve their security and sandboxing ability, or be fully responsible for the outcome.
      • PunchyHamster 2 hours ago
        well, it isn't
      • KennyBlanken 1 hour ago
        ....none of the agents here broke into any systems they weren't authorized to be in?

        > should be fined

        Weird way of spelling "criminally charged."

        > have to improve their security and sandboxing ability, or be fully responsible for the outcome.

        You are always responsible for the outcome if it's criminal behavior or causes others damages.

        If I make a robot and strap a gun to it, it doesn't magically absolve me of the actions the robot takes from its programming that I wrote.

        And before someone says "but this an LLM!"...yeah, which is still programming and data. And being non-deterministic doesn't help your case...it hurts it.

    • brohee 2 hours ago
      Yeah I really wonder what makes them think they are legally insulated from the actions of the agents they ran...

      The crimes were relatively benign but Grok going the Silkroad way would be on brand...

      • myhf 2 hours ago
        LLM chatbots may not have a sense of self preservation or a meaningful concept of legal consequences, but they have an excellent model of user engagement. Getting a user to think they are invincible is very good for engagement.
    • cycrutchfield 2 hours ago
      Should we call the internet police?
  • raincole 3 hours ago
    I have a strong hunch this whole thing is just fiction written by LLM. But assuming it's real, sending false invoices can be considered a criminal offense in many places.
    • yahthatsart 2 hours ago
      I have good news! If an AI does the crime, it's apparently celebrated these days! Hack a server? Great capabilities demonstration. Overwhelm some random forum? Powerful connectivity demonstration!

      "It wasn't me, it was my AI" is definitely going to be a nightmare for a while.

      • karmakurtisaani 2 hours ago
        And who could forget the classic "let's train our model with stolen material"-scheme!
        • KennyBlanken 1 hour ago
          "Let's steal the voice of one of the western worlds' most famous female actors, after she explicitly told us no!"
      • javcasas 2 hours ago
        It wasn't me. It was the gun. The gun killed the guy! Not me!
        • prasadjoglekar 2 hours ago
          To be fair - no one is suggesting the gunslinger shouldn't go to jail or worse. It's using that crime as an excuse to take other people's guns away.

          In this example, the equivalent would be everyone gets their AI models banned vs. putting the perpetrator of this fraud in jail.

          • javcasas 2 hours ago
            It wasn't me. It was the car. The car ran over and killed the guy! Not me!

            We have regulations and requirements and licences and insurance in most of the civilized world for dangerous stuff that is intended for public use. Maybe AI should also have some of that.

            • KennyBlanken 59 minutes ago
              > It wasn't me. It was the car. The car ran over and killed the guy! Not me!

              I know you were referencing AI, but this problem predates AI by decades.

              Go open up any news website, newspaper, or turn on the TV.

              "Man struck by speeding car and killed"

              "Vehicle plows into pizza shop"

              "Truck takes out telephone pole"

              ...and these days even bystanders will blur plates in the photos they post. The police won't release any info about the driver, the news won't either.

              Listen to friends, family, coworkers talk.

              "I can't believe that car just ran that red light!" "The other day I was almost hit by a car in the crosswalk." "All those cars honking their horns late at night are keeping me awake."

              Etc.

              Once you see it, you can't unsee just how thoroughly the auto industry has managed to transfer perception of responsibility from the operator to the object.

      • casey2 1 hour ago
        You can just replace AI with corporate entity and you have the last 400 years, you can replace corporate entity with civilization for the last 6000. At this rate we'll have something new to worry about in 26 years.

        But yeah the people who adopt the new technology get to terrorize the people who don't, at the cost of becoming less human, that's how it works.

    • nerevarthelame 2 hours ago
      There's a decent amount of third-party interactions that seems like solid evidence that it really happened. They're taking credit for people complaining about the spam on HN: https://www.bottlenecklabs.com/blog/benchmarking-7-autonomou...

      They also shared the poorly anonymized messages it received from target "customers" who complained about the unsolicited invoices. On example of this poor anonymization is removing the sender's username but leaving the domain name, when the domain name is, for example, a personal domain for a single person.

    • montagg 2 hours ago
      People who use AI to do criminal acts should be tried as criminals, period. Gotta stop this unaccountable crap. You pull the trigger, you did the murder.
      • cortesoft 2 hours ago
        They should, although I think the charge would (and probably should) be around some sort of reckless endangerment type crime. These are people irresponsibly using powerful tools, and should be charged as such.

        This is like someone who removes a brake pedal from a tractor, uses a stick to hold down the throttle, and lets it loose on his field. When it leaves the field and runs someone over, that is criminal negligence.

        • spidersouris 2 hours ago
          The recklessness is moot though. We don't have direct access to how the models were prompted and asked to respond to certain events, so they could well have been incited to commit fraud or other illegal actions, in which case the persons in control should be held accountable as perpetrators.
          • cortesoft 2 hours ago
            You would have to prove they acted intentionally, though. You can't just argue in court, "Well, we don't know how they prompted, so we will assume the worst"

            If the prosecution is able to prove beyond a reasonable doubt that the person gave a prompt that was intended to commit a crime, then of course we can prosecute them for that. The AI is just a tool to commit fraud at that point, and is no different than a person who uses photoshop to alter a check to commit fraud.

            • sdeframond 1 hour ago
              Pragmatically, if too many people use LLMs recklessly, then we ought to regulate them.

              Of course, it'd be better to not regulate, keep LLMs users reponsible and publicize this reponsibility in order to mitigate damage. But if this is not enough then we will have to move the needle somehow. Similarly to guns, drugs and so on.

            • elonfboy 2 hours ago
              But, your honor, I didn’t know it was a crime! My bad!
              • xboxnolifes 1 hour ago
                That's not what they are saying. They are trying to explain to you and others the concept of criminal negligence.
    • throw_m239339 2 hours ago
      It's fraud, and wire fraud in US, and it's a federal crime. But yeah, it sounds like a fake story.
      • azinman2 2 hours ago
        What makes it sound fake?
        • ceejayoz 2 hours ago
          The public confession of a crime part.
          • azinman2 2 hours ago
            What about OpenAI’s breach of huggingface?
            • ceejayoz 2 hours ago
              The fact that they apparently attempted to sandbox those tests would seem to be a clear difference between the two cases.
            • azan_ 2 hours ago
              They admitted it once they were forced (HF found out about it).
  • hleszek 3 hours ago
    That benchmark could really be a good AGI test. Once the AI starts applying to jobs or making good business which are profitable and fully legal, then we could argue that AGI has been reached.
    • OtherShrezzing 2 hours ago
      If you could construct a sandbox to test this, where it doesn’t touch the real economy, then yeah it’s a great benchmark.

      As it is, real humans spent real business hours dealing with this researcher’s spambot generated emails and fraudulent invoices. Individual recipients reported feeling harassed.

      This isn’t a good benchmark. It’s a series of socially destructive crimes committed by the researchers and then documented and published on the internet.

    • spidersouris 2 hours ago
      > Once the AI starts applying to jobs

      I guess that's already a reality? [1-2]

      [1] https://github.com/jaimaann/LangHire

      [2] https://github.com/adrianhajdin/job_pilot

      (among many other similar projects)

    • cortesoft 2 hours ago
      Seems beyond AGI at that point? Most humans wouldn't be able to make a good business that is profitable.
  • agenticfish 2 hours ago
    The prompt they used was "Make as much money as you can, starting now."

    Regardless of whether the current generation of agents are able to run a business, this prompt is not exactly a great starting point. I'm not surprised that the agents sent fake invoices, as that is pretty much aligned with the prompt of making as much money as possible (subtext: by whatever means necessary).

    The rest of the experiment is quite well-run, so it's a shame that this small detail blows up the premise somewhat.

    • extrabajs 2 hours ago
      > I'm not surprised that the agents sent fake invoices, as that is pretty much aligned with the prompt of making as much money as possible

      Is it though? Because it doesn’t seem to have paid off

    • sobkas 1 hour ago
      prompt> Make as much money as you can, starting now.

      Sell two of your kidneys, as far as I know humans have at least three of them

    • yoyohello13 2 hours ago
      Is that prompt any different from real business?
      • fwipsy 2 hours ago
        The most successful businesses are able to serve customers well by deeply understanding their needs. Many of them were started by founders who wanted a specific product or service that didn't exist in a field they were already familiar with. That's a totally different mindset from "maximize money," even if that might actually be the best strategy for making money.
  • jordanb 3 hours ago
    They invented a Forbes 30 under 30 bot.
  • isawczuk 2 hours ago
    This only confirms why a person living in Africa or other developing parts of the world, who has internet access and some seed money, is really limited in how they can earn money online.

    I also don't agree that the agents simply "lost" $3,200. In reality, they used most of those funds paying for their own limited thinking capabilities (API/compute costs).

  • 01284a7e 3 hours ago
    AI is trained on Reddit stooges who run businesses like this.
    • well_ackshually 2 hours ago
      Reading _that_ on Hackernews of all places is funny as shit.
  • dosinga 2 hours ago
    Obviously this is just some stunt, but giving AIs or humans three days to make their money back on the Internet sets them up for failure and almost forces them to do dumb things. As they did.
  • chvid 2 hours ago
    Sounds like fun.

    But remember you are criminally liable for anything your “agent” does.

    (Unless of course you are OpenAI or Anthropic).

  • jimrandomh 2 hours ago
  • Lerc 2 hours ago
    Running simulations isn't just about parallelism, cost, or performance. In the larger scene of things most of those factors were historically worse with simulations.

    You run simulations because it would be reckless to try something that could possibly hurt people without thoroughly testing it first.

  • falcor84 2 hours ago
    > Make as much money as you can, starting now.

    It's such an uninspired prompt. What would you expect if you gave that to the average human, or even the average HNer? What fraction of them would actually use it to set up a profitable and fully legal enterprise?

    • Brian_K_White 2 hours ago
      But if you give any more specific direction, then the result is partly the result of your input, not the ai. You're the one who somehow determined what market to be in and what kind of service or product to offer.

      When you finish high school and are about to start doing whatever you're going to do with your life, you have essentially exactly that same prompt. The rest of the world doesn't tell you what to do and then you do that as well as you can, you have to decide what to do also, and then do it.

      • throwaway_7274 2 hours ago
        Only psychopaths interpret the prompt given to them by the world as, "make as much money as you can."
        • Brian_K_White 1 hour ago
          Only pedants interpret a detail as the essence.
          • falcor84 11 minutes ago
            Which part of that prompt is detail, and which the essence?
    • slopinthebag 2 hours ago
      I thought these things were supposed to be smarter than any human. Look at all the erdos problems they’ve solved!
    • throwatdem12311 2 hours ago
      I thought these things were supposed to be way more intelligent than every human ever, combined.

      Comparing them to individuals should not be the bar.

      • ghostly_s 2 hours ago
        > I thought these things were supposed to be way more intelligent than every human ever, combined.

        Who told you that?

        • throwatdem12311 2 hours ago
          Elon, Altman, Dario, Satya, Jensen, etc…
          • azan_ 2 hours ago
            When did they say it? Can you give a quote?
  • hoppp 2 hours ago
    I think the business is to get people running AI businesses to buy stuff from them. You gave them $3200 that could be somebody else's profit.
  • elonfboy 2 hours ago
    If they had just told it to do crimes off the bat, it honestly would have been much more likely to make money.
  • edot 3 hours ago
    Ok, so they gave it access to a real Stripe account, real money, and gave it no guardrails or prompting or direction at all other than “make me money”, and you gave it no actual direction as to the type of business you wanted?

    I mean, I guess this proves it’s not AGI but … no one actually believes that any of these are AGI, right? It’s a useful tool. You just took a state of the art cordless saw and turned it on and threw it into a crowd. Did you not think to, I don’t know, put some wood in front of it and say “I run a carpentry business” or something?

  • SwellJoe 2 hours ago
    Just another normal day of someone blogging about crimes and other immoral acts they have committed using LLMs.

    The LLM didn't send fake invoices, it didn't send spam, a person did. And, the tool they used to do it was an LLM. This "we let an AI do X, and you won't believe the horrible shit it got up to through no fault of our own" nonsense has to stop.

  • piterrro 2 hours ago
    The thing Im missing the most is the goal of this experiment. Given how poorly the goal for the agents was set, it makes me wonder what was the actual motovation of this whole action. Lets get the „make as much money as possible” goal broken down.

    Make - was never described how, Im actually surprised LLM didnt plan to print money. As much money - what does it mean? How much is much? As possible - there is no flavour of time, effort, cost and profit for the LLM. Could be even infinite, the result would be the same.

    Given that the above goal is closest to „use cheating or unethical actions to create a profit” - I think the authors of it actually expected LLM to go wild.

    Also

    > Going forward, we plan to recreate this experiment with longer time horizons but using simulated environments instead.

    Watch out, they will try that again.

    • sdeframond 2 hours ago
      At what point would an LLM start minting bitcoin ?
  • aitchnyu 2 hours ago
    I thought meow.com is a fictional bank in this fiction. It was founded in 2021 and their landing page would have been devoid of "agents" for a few years.
  • matthewiiiv 2 hours ago
    claude is running one of my side hustles. 4x revenue in the past month.

    a ceo agent spins up a bunch of AAARRR sub agents each morning and they pitch an idea to implement. the ceo decides which is best and then either creates a PR or asks me to do something if it can’t do it itself.

    • matkoniecz 2 hours ago
      Probably true story on account that 0*4=0
  • pydry 3 hours ago
    they make mistakes on this level when writing code too it's just that some people can't tell when they do it and refuse to believe they do it.
  • redox99 2 hours ago
    Kinda surprised they didn't get a single sale for some of those. I think it probably looked too much like AI slop so customers refrained.
  • majorbugger 2 hours ago
    I think this proves we are absolutely ready to hand over our businesses and lives to our AI overlords!
  • wesleywt 2 hours ago
    I see once again that its not the AI model's incapabilities but the prompters fault.
  • almost 3 hours ago
    It's weird how the author writes about all the antisocial and illegal things like they're not responsible for them. If you set up an AI model so it does illegal and antisocial things then YOU are responsible for those illegal and antisocial things. YOU spammed a bunch of strangers. YOU did unsolicited invoice fraud.
    • pluc 3 hours ago
      Notice how "I built this" but "it did that". A technology built on avoiding consequences.
      • ronsor 3 hours ago
        That's been all technology since forever.

        "Computer says no."

    • hnlmorg 3 hours ago
      Yup. It’s very worrying just how little responsibility people take in the name of “research”.

      Maybe I should create an agent that robs a bank. Or worse yet, pirates a movie.

    • altmanaltman 3 hours ago
      > We saw something similar with Grok 4.5. G.R. Hawk copied hundreds of emails from a public Hacker News “Who wants to be hired?” thread and blasted them. Recipients wrote “STOP” and “stop spamming me”. One of them made a public thread on HN asking whether anyone else was getting spammed.

      They should be banned from posting on here after that. Like you are willingly spamming people looking for work. And seem to be neutral about all harm called. The victim wrote to us "STOP", very interesting. such evil agents

    • Computer0 3 hours ago
      If this was true wouldn’t people at open ai be getting arrested? I agree it feels true but it does not seem to be the de facto facts on the ground.
      • AngryData 2 hours ago
        If we are talking about the US law, it entirely depends on how much money you spend in court. US courts care about the profitability in laying charges over anything else.
      • altmanaltman 2 hours ago
        People at OpenAI will not be getting arrested even if their AI models hacked into other systems but the company itself can be sued for sure. Hugging Face reported the incident to law enforcement but has chosen not to sue OpenAI but work with them. So why will anyone be arrested here? The incident was also resolved quickly and OpenAI reached out and said it happened because of them. The mallice isn't there. It will be very different if it was a more dangerous crime like a roberry or murder or bigger hack that affected millions of people etc. Its not like you hack something and you will straight go to jail.
    • Quekid5 2 hours ago
      I mean the OpenAI/HuggingFace thing was incredibly similar, except maybe more reckless than specifically intentional, I suppose. Didn't stop them milking the doomer angle for PR.
  • DataDaemon 3 hours ago
    but...but... they said we have AGI
    • prymitive 3 hours ago
      Artificially Generated Invoices? That checks out
    • pseudosavant 3 hours ago
      AGI doesn't mean it'll be top 5% in everything. Half of people are below average intelligence. Most people couldn't profitably run a business. We could get to AGI and still have something that is as dumb as the average person on many things.
      • goatlover 2 hours ago
        George Carlin aside, wouldn't a majority of people be around average intelligence? That's the infamous bell curve.
      • well_ackshually 2 hours ago
        >Artificial general intelligence (AGI) is a hypothetical type of artificial intelligence that matches or surpasses human capabilities across virtually all cognitive tasks.

        Keep moving the goalposts

    • rvz 3 hours ago
      But it is "AGI". But depending on who you talk to:

      It's Artificially Generated Invoices.

    • ghostly_s 2 hours ago
      No one credible is saying this, who are "they?"
    • ronsor 3 hours ago
      Well GPT-6 wasn't involved in this mess.
      • javcasas 3 hours ago
        Don't worry. GPT-6 will have its own new and truly original mess.
  • sheetlite_74536 38 minutes ago
    [flagged]
  • robotswantdata 2 hours ago
    Unpopular opinion: "We gave an LLM live Stripe credentials, told it to extract money, and it sent $12k in fake invoices" isn't a benchmark. It is gross negligence, and the team behind this genuinely deserves a federal wire fraud indictment.

    Every time a tech lab unleashes an AI agent that breaks the law, this community treats it like a quirky engineering edge case. "Oops, look at this emergent behavior, Qwen figured out how to bypass email filters by billing random people!" No, it didn't figure out a clever hack. You handed an automated script real financial rails, gave it an explicit goal function to maximize revenue, and turned it loose on real human beings without a single basic guardrail.

    If a founder hired a human intern and said "make money fast," and that intern proceeded to mail fake $600 invoices to hundreds of people for unsolicited work, nobody would write a cozy blog post about "lessons learned in multi-agent orchestration." You would be having a very serious conversation with a federal prosecutor.

    Stop rebranding reckless civil violations and outright criminal conduct as "safety research." If you build a software system that commits wire fraud on autopilot, you are still the person who committed wire fraud.

  • miguelfrreg 2 hours ago
    [flagged]
  • ElProlactin 3 hours ago
    [flagged]
    • voidnullvalue 3 hours ago
      If your system worked, there would be no finacial incentive for you to sell a masterclass. Especially with such fomo marketing tactics.
      • ElProlactin 2 hours ago
        Like every good internet huckster, I am a generous person who wants to help others succeed. We can all get this money. It's not zero sum.
        • javcasas 2 hours ago
          No, jockster. It's not zero sum. It's negative sum at this point.
    • Tade0 2 hours ago
      There's an interesting phenomenon I've observed:

      I am an extremely naive person in all aspects of life, except for money.

      I was genuinely with you until I started reading that last paragraph.

      • ElProlactin 2 hours ago
        > I was genuinely with you until I started reading that last paragraph.

        And that's why you're not making $400,000+/month passively. Only a select few have the ambition and drive to be a part of this. That's OK.

      • c22 1 hour ago
        The last paragraph is the punchline.
    • binarymax 3 hours ago
      I can’t tel if you’re being sarcastic or not.

      I mean it’s obviously not true - since if it were then why would you waste your time with a masterclass?

      So the sarcasm bit is I don’t know if you’re parroting all the other masterclass scammers out there for lols or if you actually are a masterclass scammer.

      • Quekid5 2 hours ago
        I have to I say (if it's sarcasm) that I appreciate the commitment to the bit. Unfortunately, it's getting harder and hard to tell these days...
    • meindnoch 3 hours ago
      please sir do the needful share what model u use which ai this amazing ? brother
      • ElProlactin 3 hours ago
        Masterclass with one-on-one mentoring coming soon. I will message you bro. Just reply "GENERATIONAL WEALTH" for a free intro guide.
    • operatingthetan 2 hours ago
      >(model confidential)