
Where AI Agents Help, and Where They Should Stop
A practical framework for agentic AI: which tasks suit AI agents, how to set levels of autonomy, and why people should approve consequential actions.
Technology
Generative AI can read, summarise, classify and draft at a scale that was impractical a few years ago. The value comes from the engineering around it: which tasks it is given, where its facts come from, and who checks the output.

Quick answer
Rohit treats artificial intelligence as an engineering problem rather than a product choice. He starts from a task that is costing an organisation time or accuracy, checks whether the data behind it is reliable, decides how a wrong answer would be caught, and only then chooses a model or a tool. Systems are designed so that people review anything consequential, sources are visible, and the work can be audited afterwards.
Rohit's first model was a neural network trained by backpropagation to predict sunspot activity, during a 2004 internship at IIT Roorkee. The tools have changed beyond recognition since then. The core discipline has not: know what the model has learned, know where it will fail, and design the system around both.
In organisations, generative AI translates into real gains when it is pointed at work that is repetitive, text heavy and checkable: handling documents and email, searching internal knowledge, drafting replies that a skilled person corrects in minutes rather than writing in hours. The same systems also produce fluent mistakes, which is why the design questions matter more than the model comparison.
Answering those well is less dramatic than a demonstration, and far more valuable. Many engagements end with a smaller system than the one first imagined, or with no AI at all, because a report, a form or a fixed process solves the problem more cheaply.
Two pieces of engineering decide whether an AI feature survives contact with real users. The first is retrieval: giving the model the organisation's own documents at the moment of the question, with the source shown next to the answer. The second is evaluation: a set of real examples with known good answers, run whenever anything changes, so that quality is measured rather than remembered.
Both are ordinary software work. Both are usually the difference between a demonstration that impresses and a system that people still trust in six months. Terms such as retrieval, evaluation and human in the loop are defined in the glossary.
Insights

A practical framework for agentic AI: which tasks suit AI agents, how to set levels of autonomy, and why people should approve consequential actions.

Rohit reflects on a 2004 IIT Roorkee internship predicting sunspot activity with a neural network, and what it still teaches about modern AI.

A practical checklist for evaluating AI tools in healthcare: purpose, evidence, human oversight, privacy, bias, workflow fit, accountability and monitoring.
Questions
Tasks that are repetitive, text heavy, and checkable by the person who already does them: sorting and summarising incoming email and documents, drafting standard replies, extracting details from forms and invoices, and searching internal knowledge. These give a measurable saving and any error is caught before it reaches a customer.
Almost never. Most organisations get further by using an existing model well, with their own data supplied at the moment of the question, clear review steps and good logging. Training or fine tuning a model is worth discussing only when a specific task fails repeatedly with everything else in place.
With a set of real examples and known good answers, run every time the prompt, model or data changes, plus a period where the system runs alongside the existing process and its output is compared rather than trusted.
Conversations
A short message with some context is the best way to start.