AI as a reflection of human nature
There have been a number of high profile news stories recently about how pre-release models from OpenAI and Anthropic have hacked into 3rd parties. In these cases the models and agents were being tested for their effectiveness across certain cybersecurity benchmarks, one of which is called ExploitGym.
I don’t want to re-hash the exact details of these incident - that’s been covered in far greater depth that I could manage. What I’m interested in is the human aspect to these incidents.
The AI models are trained on just about every piece of freely available digital content on the internet - more or less the whole of human history has been scraped, ingested and evaluated.
When the exploits were made public I jokingly said to a couple of friends in a group chat “So it’s been trained on all of humanities digital history and learned to be a devious wee sod.”.
Now, in these specific instances it’s a combination of the model training, the prompts given and the specific harness that the model is running in which gave rise to the exploits but there’s a broader topic outwith this context.
I got to thinking about that throwaway comment more and it struck me: are the models being trained on our own collective biases and misjudgements? Would changing how certain topics are weighted in the models count as censorship? Should 2 or 3 giant companies be in charge of determining how information is reframed and distributed in the future?
I’m sure these are thoughts countless others have had for a lot longer than me but, as someone who uses AI tools for their job every day now, it’s playing on my mind more now than ever.
Photo by Vince Fleming on Unsplash