There have been many conversations surrounding AI and labour. Employers, especially those who own AI models, may frame them as an alternative to human labour, reducing the need to hire employees. Some companies promise an end to hiring humans entirely in the future. Analysis of how AI produces content suggests it may not be as independent of human labour as the marketing suggests.
How does AI produce?
AI reproduces patterns taken from the work of real people. This is why it succeeds at converting photos into certain art styles or writing cover letters. These tasks tend to follow observed patterns or even specific templates in the case of the latter.
Strict patterns help remove some of the guesswork, but they never entirely remove it, since Large Language models (LLMs) observe patterns in formatting and language instead of the actual methods that would be used to obtain the answer by a human or a purpose-built algorithm.
Some examples of this include the failure in some instances to count the number of letters in a word or perform other basic arithmetic tasks that a beginner writing a few lines of code could solve easily. In order to verify or improve output, these models require human data and human labour.
Obtaining Data
Building and maintaining an LLM requires training data. Training data can be sold directly from social media sites to the company in charge of the LLM (if they aren’t owned by the same company already), scraped by crawlers, or even volunteered by users using chat bots. The collection of data serves both to contribute to the pool of information available for the model to repurpose and to build a comprehensive user profile of what it can use to entice the user to keep using it.
The chat bots in particular will emulate emotional intimacy to encourage users to provide more detailed information about themselves. This can include medical information, sexuality, contact details etc. This allows the model to improve it’s methods of extracting money and information from the user and from advertisers who want the model to promote its product.
The repurposed content, from social media and elsewhere, is divorced from any previous contributors or context and fed to users as original content. This is particularly evident in cases where someone has asked for a source of a claim and received a “hallucination”. It takes work that humans have produced, or their bodies, and produces something in its shape without understanding the output.
Additionally, this data can be used to surveil users when sold to governments and employers. The data can be used to monitor for particular opinions or behaviours, tying those to more public identifiers.
In this way LLMs harvest and alienate internet users from their labour and from each other. Users are charged for the privilege of receiving information and work provided by other users for free.
Obtaining Labour
Human labour is just as crucial a commodity for LLMs, on every layer of operation. From the physical layer of materials and infrastructure to the information layer of moderation and analysis.
On the layer of physical infrastructure, LLMs rely on land, energy and manufacturing. The mineral extraction in places such as DR Congo, where Western funded militias drive people into desperate poverty in order to gain cheap materials for processors and chips. This is continuation of western capitalist powers taking both labour and land and using one to reinforce the other.
The information layer relies on both knowing and unknowing labour. The knowing labour includes content moderation and data labelling. These jobs, while not necessarily as physically exploitative as mining or construction, frequently require long hours for low pay, low job security and at a psychological cost to the labourer. This can include being subjected to images and videos featuring real graphic violence and CSAM, particularly in the case of content moderation. This work, like maintaining the physical infrastructure, is typically done by people in the imperial periphery, which is what allows the owners of AI models to extract so much while paying so little.
The unknowing labour is collected more indirectly. Internet users have, without informed consent, been training data for AI as well. ReCAPTCHA, the test implemented on many websites to “check that you’re not a robot”, was sold to Google for AI training. Millions of users were forced to contribute training data so that AI models could parse strange looking text and identify objects.
Websites with targeted algorithms like Spotify or Tinder frequently sell data to third parties. Previously for use by targeted advertisers, now also used for training AI models, if those models don’t manage the algorithms directly. Information about shopping habits, conversations between users or even users voice data, is made accessible to these models.
The analysis and pattern recognition is divorced from the labourers performing it. It is sliced up and profited from by the capitalists who own the models, while the model would not function without human labour.
What can be done?
The exploitation described above is not unique to AI. The abuses of privacy or workers were features of a system dependent on the internet prior to AI. AI models have made an already extractive system more extractive.
A technology dependent on human labour is vulnerable to human organising. This will involve organising the tech workers, physical and data labourers, as well as finding and building alternatives to the technologies required in our daily lives. Currently the tech industry in Aotearoa is largely deunionised. To address this, a campaign of unionisation by tech workers and a broader intersectional struggle is required.

