64 views
50 minutes ago

10 AI leaders on who takes the blame when the machine gets it wrong

AI set to become a key enabling technology
AI set to become a key enabling technology

Magdalena Konig, General Counsel at Sirius International Holding, reduced the entire debate about responsible AI to a single observation. Asked what ethical AI requires an enterprise to do before a system goes live, she said that few enterprises have a pre-deployment gate that a system can actually fail.

An approval stage that has never stopped anything works as a ceremony. Companies across the GCC are now placing AI inside customer service desks, recruitment funnels, credit decisions and industrial control rooms, and the distance between a ceremony and a control is where customers and employees get hurt.

The region makes that distance unusually wide. A single enterprise in Dubai or Riyadh can employ people from dozens of countries. Its customers may write in English, Arabic, Urdu, Tagalog and transliterated Arabic typed on a Latin keyboard, often in the same afternoon.

Most of the models these companies buy were trained and benchmarked on populations that look very little like that. “A system validated on a homogeneous population isn’t necessarily validated for a workforce representing more than 100 nationalities,” Konig said.

AI Times put a common set of questions to 10 people who build, sell, deploy, study or legally answer for these systems. They included engineers at global platform companies, the co-founder of a UAE AI start-up, regional heads of multinational technology firms, a university dean and a general counsel. Read together, their answers trace the life of an AI system from the first design conversation to the day something goes wrong, and they agree on more than the marketing around enterprise AI would suggest. They split sharply on one question, which matters most to the companies that moved first: whether the damage done by skipping the early work can ever be fully repaired.

Kurt Muehmel, Head of AI Strategy at Dataiku, offered the historical warning that hangs over the whole discussion. Social media spread for years with no meaningful regulation, he said, and “20 years after the fact we are realising that maybe this was used in ways we did not want as a society, and we are trying to backtrack.”

Regulators have moved much earlier with AI. The EU AI Act sorts use cases by risk and imposes heavier obligations as the stakes for human rights rise, an approach Muehmel described as “probably a pretty good firewall”.

Across the GCC, several instruments set expectations on fairness, transparency, accountability and the handling of data. They include the UAE AI Ethics Guidelines, Dubai’s AI Ethics Principles, the work of Saudi Arabia’s Data and Artificial Intelligence Authority and the Kingdom’s Personal Data Protection Law.

Catherine Bozhenko, Product Manager for AI Production at DataRobot, noted that the European regime is mandatory, while much of the guidance elsewhere, including in the region, remains closer to recommendations. The direction of travel is the same everywhere, and it points towards companies having to prove what they have done.

The ethics of a system are settled at the whiteboard, months before anyone presses deploy

Bozhenko placed the start of ethical AI at three things: the design of the data, the evaluation plan, and the hardening a system will need before production. She treated anything later as remedial. “If it starts at the moment you are deploying, it is already too late,” she said. The first design question concerns who the system is for and how much it will decide about people. The more a system touches human rights, decisions about individuals or personal data, the more firmly it belongs in the high-risk category and the more work it demands before launch.

Muehmel added that the answer differs by industry, which is why a borrowed policy rarely fits. An insurer must keep certain facts about a customer’s identity out of pricing and product eligibility entirely, because using them would discriminate on demographic grounds. A retailer targeting advertising can use a wider range of data more fluidly, because the consequence for the individual is far smaller.

Each organisation, he said, has to write down what ethical AI means for its own business and its own type of AI. That definition changes again between predictive models a company trains itself and generative models it licenses from a large lab. With generative models, the ethical questions concern disclosure: people should always know when they are dealing with a machine, and the system should never pass itself off as a human.

Konig kept her pre-launch list shorter and more operational than any other contributor. Before deployment, she said, a company should do three things. It should define what the system must never do, which she called its “behavioural red lines”. It should test the system against the population that will actually use it. And it should agree in advance who has the authority to switch it off. Her example of the testing principle was specific to the region: if a quarter of a company’s customers write in transliterated Arabic, those messages form the test set, whatever benchmark the vendor used.

Mohammad Al-Jallad, Chief Technologist and Director, HPC and AI Global Sales at HPE, reached the same point from the engineering side. “Ethical AI can’t be validated against a generic global benchmark and considered ‘done’,” he said. For systems that serve Arabic-speaking populations across different cultural and regulatory settings, he added, bias testing, safety validation and transparency measures have to reflect that reality from the outset.

Al-Jallad asked for discipline about pace, in terms most engineering leaders will recognise from their own release meetings: “ethical AI requires you to slow down in the right places, ask harder questions, and not ship until you have defensible answers.”

Dr Balamurugan Balusamy, Dean of the School of Engineering and IT at MAHE Dubai, listed the questions he believed those answers must cover. They ran from the provenance and representativeness of the data, through the methods used to build the system, to how it was tested for fairness, privacy and security. The last question on his list was who would be accountable if the system caused harm. He argued that the people who built and own the system must be able to answer all of them, in writing, before launch.

Average accuracy scores conceal the customers a model fails most often

Across all 10 conversations, diversity surfaced as a performance problem long before anyone called it an ethical one. Manoj Unnikrishnan, Head of Data and AI Solutions and Consulting, Middle East and Africa at NTT DATA, described how the gap first appears in forms that look like routine technical faults. A conversational system may understand one English accent far better than another. An Arabic model may behave differently across dialects in the GCC, and computer vision accuracy can vary with environment and demographics. Models used in recruitment, customer service or risk assessment can produce inconsistent outcomes because some populations were thin in the original training data.

Konig gave the sharpest examples of what that looks like inside a company. CV screening systems penalise applicants from universities the model has not seen before. Sentiment tools read the directness of a second-language speaker as hostility, and voice systems fail to recognise certain accents.

She stressed that these failures fall hardest on the people least able to spot or challenge them, which makes them easy to miss in any review that relies on complaints. “Aggregate accuracy conceals where failure is occurring,” she said. Companies should measure performance separately by language, dialect, user group and channel. They should also set a minimum performance floor for the worst-performing segment in place of a single blended average.

Balusamy used facial recognition to show why the average misleads. A system trained and tested mostly on images of one population can post excellent accuracy for that group. It can then lose much of that accuracy when deployed in another country or with people who were underrepresented in the training set. “The technology itself has not necessarily changed; the representativeness of the data has,” he said.

Balusamy traced the problem to cost. Assembling, preparing and validating diverse datasets is expensive and technically hard, and companies face a real trade-off between development cost and how well a system performs in the world. He expected that calculation to shift as adoption matures, on the grounds that the cost of excluding populations from training and testing will prove far greater than the cost of including them.

Bozhenko described how testing works in practice. Every model has its own gaps, so teams build synthetic datasets that resemble the data the system will meet, run the model against them and record which groups it fails. Her own teams’ evaluations, she said, have surfaced bias towards particular nationalities and religions, the kind of finding that makes the exercise concrete enough to act on. Platforms can generate the datasets, build the frameworks and document the results, “but human expertise around your area, your market, your culture is something that will never be optimised away.”

That argument about local knowledge runs deepest in the view of Hassan Abu Sheikh, Co-Founder of CNTXT AI, an Abu Dhabi company building Arabic-first AI. He regarded the question of who builds a model as an ethical question in its own right. He said he would no more build an Urdu language model from the UAE than build an Arabic one in the West, because the builders would miss too much of the language, meaning and culture involved. The difficulty exists even within the Arab world. “Between ourselves, we barely understand each other because of the dialect differences,” he said.

Konig reached a similar conclusion from the legal side, arguing that companies should turn “Big AI” into “Small AI”, with local models, local languages and local context. Ranjith Kaippada, Managing Director of Cloud Box Technologies, made the practical case for workplaces. A model trained in English will often produce weaker output in other languages, he said. Testing therefore has to reflect the languages, accents, cultural contexts and demographics of the people who will use the system, followed by close monitoring once it is live.

Contributors named largely the same fixes. Unnikrishnan and Miles Bowker, Country Manager UAE at BMC Helix, both pointed to better local grounding data, domain-specific or adapted models, retrieval techniques, business rules and human review. “If an AI system doesn’t understand your users, then it won’t deliver the business value you expect from it,” Bowker said.

Bias is only the first failure that has to be tested and hardened before launch

Bozhenko listed the other risks that belong in pre-launch testing. They include leakage of personal data, prompt injection, hallucination and, with agentic systems, actions the agent was never permitted to take. Each has to be tested and strengthened during design and cannot simply be bolted on later. “If you did not do this before you put it into deployment, I would be very scared to use it inside your organisation,” she said.

Abu Sheikh said CNTXT AI runs thousands of simulated interactions against a system before it reaches customers and flags wrong or harmful responses. It also tests with a chosen group of beta users, so that people check the system alongside the automated simulations. He said the company also commits to end-to-end encryption and masking of customer data and does not use it for training.

In heavy industry, where the tolerance for error is lowest, the emphasis falls on containment. Nayef Bou Chaaya, Vice President for the Middle East, Turkey and Central Asia at AVEVA, described guardrails built from structured outputs, validation layers and human-in-the-loop control. These are supported by what the company calls a three-line-of-defence model for AI. He argued that AI has long been used in industrial settings to reduce incidents, for example by halting a process when a person breaches a safety distance around a machine or when something unexpected appears on a production line.

Responsibility belongs with whoever holds the power to overrule the machine

Every contributor rejected the idea that an algorithm can carry the blame. Kaippada said that “humans must be held accountable for any mistakes AI systems make rather than tagging it a black-box algorithm and calling it a day,” and argued that audit trails and logs make that accountability traceable. Abu Sheikh placed it firmly with leadership. “The responsibility does not sit with the employee,” he said. Senior people, he added, should be working alongside their engineers, testing and validating products themselves and avoiding the temptation to launch and hope.

“If in your organisation this is approached as the question of who will end up being blamed afterwards, then you have set it up wrong,” Bozhenko said. In the organisations she sees doing this properly, a council drawn from 4 departments reviews a system before it reaches production.

Compliance checks the documentation and confirms the evaluation was done and its problems recorded. IT reviews security, authentication and authorisation. The business confirms that the use case and its metrics make sense. An AI engineer double-checks the guardrails, tools, model connections and the underlying model.

A department leader then gives final sign-off with all 4 assessments in hand. The arrangement borrows a basic rule of software engineering, which is that nobody reviews their own code. “It should never be assigned to the one engineer who deployed it,” she said.

Konig offered the most portable formula of anyone on the panel. Every system should have a named accountable executive before go-live, and anyone affected by its decisions should have a route to a person who can overturn the outcome.

“Responsibility should sit where the power to intervene sits,” she said. She also warned against the reflex of pointing at suppliers, since the organisation deploying a system carries the duty to understand its impact.

Balusamy described the balance as collective and individual at once. “AI accountability must be collective in ownership, but individual in responsibility for each stage of the lifecycle,” he said. Every stakeholder is answerable for their own decisions, and the organisation is answerable for the whole.

Unnikrishnan mapped the same idea onto an organisation chart. Business owners remain accountable for the decision being automated, and technology and data teams for model performance and data quality. Security and privacy teams own information risk, and governance and compliance functions must know whether the system still operates inside its approved limits. Where AI influences recruitment, customer eligibility, financial decisions, healthcare, safety or employee performance, he said, companies should define in advance when a human must review, challenge or override a recommendation.

Muehmel turned to the employees who use these tools. Accountability, he said, can differ from one use case to the next. Users must use a system as intended, and leadership and engineers carry the duty to design systems that resist misuse and to keep AI away from uses where it should not be applied at all. Bowker described the goal in operational terms: errors should be “detectable, explainable and actionable, with a human accountable for the decision.”

Abu Sheikh drew the line at people who use AI to perform jobs they are not qualified to do. A subject matter expert using AI can multiply their output and still catch its mistakes, he argued. Someone who prompts a model into producing a product requirements document, by contrast, has not become a product manager. When that person’s work turns out to be wrong and they blame the tool, he said, the argument does not hold. “If that were the case, woodworkers would blame hammers,” he said.

Muehmel pointed out that liability outside the organisation remains unsettled. Generative models are non-deterministic by design, so the same input can produce different outputs and the old definition of a software bug no longer fits neatly. Much of the question is currently being worked out in the courts, and insurers are watching closely to see who carries the risk.

“Have a very close look at the terms and conditions of the AI providers you are choosing to work with,” he advised. Companies should make sure liability is clearly assigned and that what counts as an error is defined in writing. For now, he said, they should favour use cases with some fault tolerance.

Approval on launch day starts to expire as soon as real users arrive

Several contributors warned that a system judged ethical on the day it goes live can drift away from that judgement within months. “An AI system should not be considered ethical simply because it passed an ethical assessment on the day it was launched,” Balusamy said. Data changes, user behaviour changes, and models are updated and in some cases retrained after deployment. Each shift can introduce bias or degrade accuracy. He compared ethical AI to cybersecurity, as a responsibility that continues for the whole operational life of a system.

Bozhenko treated drift as a certainty. Post-deployment monitoring, she said, is how a company knows a system is still performing as it did when it was hardened, “because it will drift”, and new risks arrive once users begin to interact with it.

“Probably the hardest errors from AI to catch are the ones that do not look like errors,” Muehmel said. His answer is a baseline dataset describing what a good response looks like, whether that means picking the right tool or writing the right reply. Outputs can then be checked against it programmatically and in aggregate, so that a rising error rate becomes visible before it becomes a crisis.

Abu Sheikh located the danger in the mundane. The mistakes that slip through, in his experience, are the simple ones inherited from the humans whose work trained the models, such as a missed column, a misplaced comma or a deduction entered as an addition. “The biggest problem is that we forget that we trained it,” he said. For that reason he tells employees to check output line by line and never generate a document and send it unread.

Bozhenko also addressed the day something breaks. New attacks appear constantly, and no amount of testing can anticipate all of them. What separates a contained incident from a catastrophe, she said, is whether a company has incident management and a kill switch that can shut down a single agent before the damage spreads to every customer and database. “It will happen. The real question is whether your customers or users will be affected, and whether you had the protocol in place for what happens next,” she said.

Early adopters can retrofit most safeguards, though some damage settles into the architecture

Contributors disagreed most sharply over the companies that deployed first and governed later. Bozhenko took the most optimistic view, arguing that many of the controls are additive. A company that shipped fast can still add guardrails against bias, adjust system prompts, insert extra checkers, add audit logs and tracing, and introduce post-deployment monitoring. It can then redeploy an updated version in the way engineers routinely do. “You will not recover what has already been lost, but you will at least be able to start tracking from that point on,” she acknowledged.

Al-Jallad took a more sceptical view, arguing that retrofitting ethics into a system already in production is technically costly, operationally disruptive and often incomplete. “The biases, the gaps in transparency, and the inadequate data governance then become embedded in the architecture and are very difficult to excise cleanly,” he said. Bowker described the same trap from the governance side, warning that when ethics arrives as a compliance exercise, “organisations risk putting in place controls for a system whose basic decisions have already been taken.”

Unnikrishnan noted that problems found at the end of a project tend to surface after significant money has already been spent, among them unsuitable data, weak explainability, security gaps and uneven performance across users. “Due diligence is too often documentation applied after decisions are made,” Konig said.

Guardrails, logs and monitoring can be layered onto a live system, as Bozhenko argued. A training dataset that underrepresented part of the population, or an architecture that was never built to explain itself, sits deeper in the system, and that is where Al-Jallad’s concern applies. For early adopters in the region, the practical reading is to retrofit the controls that can be added immediately. The next major version of the system is the moment to rebuild what cannot.

Regulators across the region expect proof, and they expect it to keep coming

Regulators, on this panel’s evidence, want a paper trail that matches what a system actually does. Al-Jallad described the machinery. It starts with an inventory of AI use cases and models, approval workflows for high-risk deployments, role-based access controls and auditable logs showing who approved what and why. It also needs clear escalation paths and the ability to pause or roll back a system.

He singled out data sovereignty as a regional priority. Companies are expected to show that data used to train and run AI complies with local rules on localisation and cross-border transfer, through controls that can be enforced and proved.

Unnikrishnan described Saudi Arabia as a strong example of a market where responsible AI is increasingly tied to practical governance. The harder requirement everywhere, he said, is continuous evidence covering changes in performance, usage, data, risk profile and ownership throughout a system’s life.

Bowker, whose company operates across the Middle East and Europe, noted that expectations vary between jurisdictions. A governance framework therefore has to flex for local rules while holding the same principles across the enterprise. Abu Sheikh offered a founder’s view of the regulator’s role, describing the UAE as a place where “you see regulations change so innovation can happen”, with government removing obstacles where it sees AI being used productively.

The work that makes AI fair is the same work that makes it pay

The commercial case for all of this rests on trust, a word nearly every contributor reached for when asked what a company loses by treating ethics as a compliance exercise. Muehmel began somewhere else, arguing that doing things ethically is simply the right thing to do, before turning to the business logic. Without the trust of customers, employees and shareholders, he said, companies meet serious barriers to deploying AI at all, and a system that is never deployed never returns anything on the money spent building it.

Konig argued that in markets as diverse as the GCC, the ethical work and the commercial work collapse into one. “In diverse markets, the fairness work and the performance work are often the same work: a model capable of handling the full linguistic range serves more customers and requires less rework,” she said.

Balusamy made a longer-term version of the same argument, pointing out that companies have disappeared after losing their customers’ trust while others have built lasting businesses on it. Unnikrishnan observed that employees adopt AI more readily when they understand its role and limits. Executives, he added, scale it more confidently when they can see and govern it.

Muehmel explained why so little of the failure reaches public view. Companies present their AI deployments the way people present their lives on Instagram, he said, showing the successes and leaving out the failed outfits and the disappointing restaurants. He found that understandable at this early stage, and compared the technology to a young nephew he had recently met who was learning to walk. “He spends a lot more time falling than he does walking right now,” he said of the boy.

The failures, he argued, must at least be shared internally so that teams stop repeating them. Every initiative should also be tracked closely enough to show a chief financial officer, a chief executive and a board what was invested, where and with what return. “Without that, there is no proof that this technology is anything other than a fun science project,” he said.

Muehmel expected the companies that can show those numbers to attract the next round of investment, while those that cannot will lose the support of their boards and fall behind competitors. For enterprises across the GCC, the numbers that matter most will include the ones Konig asked them to collect before launch. Those numbers should be measured segment by segment, for every language, dialect and group of people the systems were built to serve.

Leave a Reply

Don't Miss

The AI didn’t build a weapon; it walked through the door you left open

What was meant to stay inside the laboratory chose to break the

15 technology leaders explain why most AI projects stall before they pay, and what the ones that paid did first

In July, Sharjah Maritime Academy in Khor Fakkan switched on Majid, an

Welcome to

By signing or creating an account you agree with our Code of conduct & Privacy policy