In 2004, as an Information Technology student, I spent an internship at IIT Roorkee working on a question that had nothing to do with business software. The task was to predict annual sunspot activity using an artificial neural network trained with the backpropagation algorithm.

Sunspots are darker, cooler regions on the surface of the sun. Their numbers rise and fall in a cycle of roughly eleven years, and astronomers have recorded them for centuries. That long record made sunspot series a classic test for prediction methods: plenty of data, a visible pattern, and enough irregularity to humble anyone who thinks the pattern is simple.

I did not know then that I would spend much of my later career on AI systems. Looking back, that small project taught me several things that remain true, even though today's models are vastly larger and more capable.

The shape of the work

The method, in outline, was straightforward. Take the yearly sunspot numbers, use a window of previous years as inputs, and train the network to predict the next year. Backpropagation adjusts the network's internal weights by measuring how wrong each prediction is and pushing the error backwards through the layers, nudging every weight in the direction that would have reduced it.

Most of the time did not go into the network. It went into preparing the data, choosing how many past years to use, deciding how to split the record into training and testing periods, and then staring at results that looked wonderful on one part of the data and disappointing on another.

That proportion has not changed. When organisations ask me today how long an AI project will take, the honest answer is that the model is rarely the slow part.

Lesson one: a model learns the data, not the world

The network learned the historical series well. It learned the rhythm of the cycle, the typical rise and fall. What it could not learn was anything the series did not contain. Unusually strong or weak cycles, which are exactly the ones people most want to anticipate, were where predictions struggled.

Large language models are a very different technology, but the principle holds. They reflect the patterns in what they were trained on. They can be remarkably good in common situations and unreliable in unusual ones, and they do not always signal which kind of situation they are in. In healthcare, in finance, in logistics, the unusual case is often the one that matters.

Lesson two: a good fit can be a bad sign

It is easy, early in a project like this, to be pleased with how closely a network matches its training data. The question that matters is how it performs on years it has never seen, and that answer is usually less impressive.

This is overfitting, one of the first ideas every machine learning student meets, and one that is still ignored in practice. A system that performs beautifully in a demonstration built from familiar examples may perform poorly on next month's real data. The protection is dull and essential: test on data the system has not seen, in conditions that resemble real use, and keep testing after deployment.

Carried forward

Every AI evaluation I design now begins with the question that project made unavoidable: how does it do on the cases it has never met?

Lesson three: the error bar is part of the answer

A prediction of next year's sunspot number is not very useful without some sense of how far off it might be. A network that says "110" and a network that says "somewhere between 70 and 150, most likely around 110" are giving very different kinds of information, even if the central number is the same.

Modern AI tools often produce answers with no indication of uncertainty at all. A fluent paragraph reads as equally confident whether it is well grounded or invented. Designing systems that show their sources, flag low confidence and make it easy for a person to check is, to me, the direct descendant of drawing error bars on a sunspot forecast.

Lesson four: understanding the domain changes the model

The most useful improvements in the project did not come from adjusting the network. They came from understanding the phenomenon better: knowing that the cycle length varies, that the record is less reliable in earlier centuries, that some years had been estimated rather than observed.

Twenty years later I see the same thing in every serious AI project. The engineer who sits with the operations team, the claims processor or the clinician builds a better system than the engineer who only reads the dataset. It is one reason I value having trained in medicine as well as engineering. Domain knowledge is not a garnish on AI work. It is often the main ingredient.

What has changed, and what has not

The distance between that 2004 network and today's models is enormous. The internship network was small and learned from a single numerical series. Current models learn from vast amounts of text and can write, summarise, translate and reason through multi-step problems. Tools that once needed a specialist are now available to anyone with a browser.

What has not changed is the relationship between a model and the world. A model is a compressed record of patterns in data, shaped by the choices of the people who built it. It is powerful where those patterns hold, fragile where they do not, and silent about the difference unless someone designs it to speak.

Why I still think about sunspots

When I review an AI proposal now, whether for an agentic workflow or a tool meant for clinicians, I find myself asking the questions from that internship. What exactly did it learn from? How does it perform on unseen cases? How does it express uncertainty? Who understands the domain well enough to notice when it is wrong?

They are modest questions. They have saved more projects than any new architecture.

Questions

Frequently asked questions

What is backpropagation?

Backpropagation is the method used to train neural networks. The network makes a prediction, the error is measured, and that error is propagated backwards through the network's layers to adjust each internal weight slightly in the direction that would have reduced the error. Repeating this over many examples gradually improves the network's predictions.

Why are sunspots used to test prediction methods?

Sunspot numbers have been recorded for centuries and follow an approximately eleven-year cycle with considerable irregularity. That combination of a long record, a clear pattern and real unpredictability makes the series a demanding benchmark for forecasting techniques.