AI & Data Science leader with 12+ years across application engineering, business analytics, machine learning, predictive modelling, fraud and risk analytics, Decision Science, GenAI and AI application development.
Builds systems that connect data, models, explanations and recommendations to real business decisions — with an architecture path toward governed Agentic AI.
Quantified outcomes are presented with their stated measurement window; methodology and attribution are available for discussion. Selected work is presented with client and employer details anonymised where required due to confidentiality and consulting engagements — full employment history is provided in the CV.
The same seven steps, whatever the technology.
The same sequence applies whether the output is a scorecard, a GenAI platform or a mobile product. Select any stage.
Seven capability areas, each with working proof behind it.
Each one has working evidence behind it.
LLM applications, natural-language analytics, diagnostics, recommendations, governed tools and AI product design.
Tool orchestration, specialist-agent patterns, human approval, guardrails, auditability and action workflows.
Predictive modelling, classification, regression, forecasting, anomaly detection, evaluation and explainable AI.
Fraud analytics, network and graph analysis, credit risk, expected-loss prioritisation, anomaly detection and financial-services decisioning.
Pricing, revenue, customer and operational decisions — connecting predictions and explanations to recommended actions.
Python, R, SQL, APIs, web applications, Dataiku, Azure, Databricks and GCP.
Problem framing, stakeholder alignment, consulting delivery, team leadership, mentoring and adoption.
Engineering to analytics to ML to Decision Science to enterprise AI.
Engineering to analytics to machine learning to Decision Science to enterprise AI. Each phase built the foundation the next one needed.
Nine layers, from the question a person asks to a governed action.
Nine layers. Select any one to see what it does and why it exists.
Stated in three explicit levels, because conflating them is the fastest way to lose credibility in a technical interview.
Intent routing between deterministic and reasoning paths, retrieval with a relevance floor and enforced refusal, SHAP-grounded diagnostics, tool-style calls into a governed semantic model, and full audit logging of prompt, retrieved context and response.
The multi-agent pattern below, running on synthetic data inside the embedded platform. It shows the orchestration and the approval gate working — it is not a production autonomous deployment, and is not presented as one.
Specialist agent decomposition, governed write actions back into source systems, cost and rate control per agent, and evaluation harnesses for multi-step task completion rather than single-turn accuracy.
The approval gate is the architecture, not a limitation of it. An agent that can act without a named human on irreversible decisions is a liability in any regulated environment.
Eight applications you can open and use right now.
Interactive working demonstrations using synthetic data. Several put you in the operator's seat — approve discount requests, underwrite applications, tune a fraud threshold and see what it costs.
Where a language model helps — and where it must not be used.
A language model is excellent at language and unreliable at arithmetic. The design question is not "how do we add an LLM" but where it is allowed to touch the answer. Deterministic analytics and a governed semantic model handle numbers; the model handles reasoning and narrative; evidence and access boundaries are enforced around both.
Leadership could see that commercial turnaround was slow. They could not see why. Answering one diagnostic question meant an analyst pulling data for two days, and the same questions recurred every month while the effort never reduced.
A dashboard answers what. It cannot answer why — attribution needs a model — and it cannot answer questions nobody built a tile for. Adding tiles made it worse: the estate grew while the questions kept arriving by email.
Leadership could see commercial turnaround was slow. Nobody could say why without two analyst-days per question.
Where to intervene in the approval chain — and whether the cause was routing, submission quality or deal mix.
Semantic model for numeric truth; attribution model for cause; LLM constrained to narrating retrieved evidence.
Retrieval hit rate, groundedness and refusal behaviour on a held-out question set. Framework below.
Diagnostic questions answered in seconds rather than days. Synthetic demonstration — outcome figures not published.
The obvious design maximises coverage: let the model answer everything, retrieving where it can. That produces a system which is impressive in a demo and untrustworthy in a meeting, because the failures are indistinguishable from the successes.
I chose the opposite. A relevance floor means questions without supporting evidence are refused rather than answered, and numeric questions bypass the model entirely. The cost is real — the system declines questions a human could answer, and users notice. The gain is that when it does answer, the answer can be checked against a cited source.
Second trade-off — latency and cost vs model quality. Routing most traffic to the deterministic path costs nothing per call and returns sub-second; the reasoning path costs a model call and several seconds. That routing decision has a direct line to the operating bill, which is why it is an architecture choice rather than a runtime one.
Deterministic systems for deterministic truth. Machine learning where prediction adds value. Language models where reasoning or synthesis adds value. Choosing wrongly is the most expensive decision in an AI programme, and it is made early.
| Metric | Test set | What it tells you |
|---|---|---|
| Retrieval hit rate | Measured | Does the document containing the answer reach the model's context? If retrieval misses, nothing downstream can recover — most bad answers are retrieval failures in a generation costume. |
| Grounded answer rate | Measured | Is every factual claim traceable to a retrieved passage? A claim with no supporting passage is a hallucination even when it happens to be correct. |
| Appropriate refusal | Measured | Held-out questions the corpus deliberately does not answer. The system should decline all of them. |
| Over-refusal | Measured | The opposite failure, and the one that quietly kills adoption — an assistant that shrugs too often is abandoned within a fortnight. |
| Unsupported answer rate | Measured | Answers produced without adequate evidence. The metric that matters most, and the hardest to move without sacrificing coverage. |
Values are computed live inside the demonstration against its synthetic corpus. No production evaluation figures are published here, because none would be defensible outside the engagement.
Failure mode. Early versions answered confidently on questions the corpus did not cover. The output was fluent, specific, and wrong — the worst possible combination, because it was indistinguishable from a good answer.
Why it happened. Retrieval always returns something. Passing the top-k results to the model regardless of their scores meant weak matches were treated as evidence.
Design change. A relevance floor below which nothing is passed to the model at all, plus an instruction to answer only from supplied context and to state plainly when it cannot.
Result. Unsupported answers largely disappeared; over-refusal appeared in their place, which is a far cheaper failure and a visible one.
Remaining limitation. Keyword retrieval still misses paraphrase — a user asking about "turnaround" will not match a document saying "cycle time". Dense embeddings, or a hybrid, are the fix at production scale, and are not implemented in this demonstration.
A second system, same principle. Retrieval scores documents against the query; only those clearing a relevance floor reach the model; every answer cites its sources. When nothing clears the floor, it refuses — and the evaluation is built around that refusal rather than around accuracy alone.
Retrieval hit rate, groundedness, appropriate refusal and over-refusal. A system that answers every question is broken, not finished — some questions have no documented answer, and the correct response is to say so.
Systems that decide: fraud, credit, pricing, forecasting.
Each row is a decision somebody is accountable for, and each has a working demonstration behind it.
| Decision | What the system does | Try it |
|---|---|---|
| Fraud | Anomaly detection with a rule layer, ranked by expected loss rather than score so investigator time follows the money. Approximately $1M in estimated fraud losses prevented over three quarters. | |
| Financial crime | Graph linkage across shared devices, instruments and destinations. Union-find isolates communities; betweenness centrality separates the coordinator from the mules. | |
| Credit risk | Points-based scorecard with weight of evidence, probability of default and reason codes. Deliberately not an ensemble — every point is traceable, so a decline can be explained. | |
| Pricing | Elasticity per segment, and the finding that changed the conversation: margin falls monotonically with discount, so only a deal you would otherwise lose justifies one. | |
| Forecasting | Holt-Winters with rolling-origin backtesting against naive baselines, prediction intervals rather than point estimates, and PSI drift monitoring that triggers retraining. |
For fraud, financial crime and credit, the system improves the speed and quality of a decision. It does not become the decision-maker. Every high-risk path routes through a person who can be named in an audit.
The threshold is where the design judgement sits: too high and losses pass through silently, too low and investigator capacity is consumed by false positives. The demonstrations let you move it and see what it costs.
Every point of threshold you lower buys recall and costs investigator hours. Both are measurable, which means the operating point is a commercial decision rather than a modelling one — and it should be made by the people who own the loss budget, not by whoever tuned the model.
Accuracy vs explainability. A gradient-boosted ensemble scores better on the credit book than a points-based scorecard. I shipped the scorecard, because a declined applicant is entitled to a substantive reason and "the model said so" is not one. That is a deliberate accuracy sacrifice for a regulatory requirement, and I would make it again.
What this costs. Perhaps two points of Gini. What it buys is a model that survives audit, that an underwriter will actually action, and that can be defended in front of a regulator without a translation layer.
Failure mode. The first model looked excellent in sample and could not beat a naive "assume no change" baseline out of sample.
Why it happened. In-sample fit rewards a model for describing noise it has already seen. On a financial series most of the variance is not forecastable, so a flexible model fits the noise and carries it forward.
Design change. Rolling-origin backtesting against naive and seasonal-naive baselines became the gate. If a model cannot beat "no change", it does not ship.
Result. The model that shipped wins clearly on seasonal, flow-driven components — and the demonstration will show it losing to the baseline under a regime shock, because that is the honest result and it is what the technique exists to surface.
Remaining limitation. No amount of tuning survives a policy shift or a devaluation. The correct response is to detect the break quickly and refit, not to claim the model anticipates it.
Fraud, risk and financial decisioning, gathered in one view.
The work most relevant to payments, banking and financial services — gathered in one view. Nothing here has been relabelled to fit; each item is financial-services work as delivered.
| Capability | What was built | Evidence |
|---|---|---|
| Fraud detection & expected-loss prioritisation | Anomaly detection with an explicit rule layer, ranking cases by expected loss rather than model score so investigator capacity follows financial consequence. ~$1M in estimated fraud losses prevented over three quarters. | |
| Graph & network analytics | Linked-entity analysis across shared devices, IP ranges, funding instruments and destinations. Union-find isolates communities; betweenness centrality separates coordinator from mules. | |
| Credit risk & PD | Points-based scorecard — weight of evidence binning, probability of default, reason codes, and full traceability of every score contribution. | |
| Anomaly detection | Applied across claims and transaction populations, with thresholds tuned against investigator capacity rather than against a headline metric. | |
| Financial forecasting | Currency forecasting supporting treasury planning, validated by rolling-origin backtesting against naive baselines, delivered with prediction intervals. | |
| Pricing & elasticity | Segment-level elasticity estimation driving discount decisions, with the margin trade-off made explicit rather than assumed. | |
| Explainable AI & decision governance | Reason codes falling out of model structure, SHAP attribution grounding diagnostics, retrieval-level permissions, and audit trails on every automated recommendation. |
Not just models — applications people actually use.
Enterprise analytics and product engineering are different disciplines. This is the second one — shipping something real, to real users, with the operational consequences that follow.
Society administration and revenue operations run on spreadsheets, WhatsApp threads and paper receipts.
Who owes what, whether it was collected, and who is accountable for chasing it.
Mobile client, API layer, authentication, application services, database and notifications — built and deployed end to end.
In production use. Adoption and operational metrics available on request rather than published.
Committee, resident and staff working from one record instead of three. Production application.
A society management application handles money. That argues for heavy controls from day one — approval workflows, segregation of duties, full audit. It also argues for shipping nothing, because a committee of volunteers will not adopt a system that makes their monthly routine harder.
What I did. Audit trail and role separation on financial actions from the first release, because those are unrecoverable if retrofitted. Everything else — reporting depth, workflow configurability — deferred until the core was being used.
Remaining limitation. The controls are proportionate to a residential society, not to a regulated financial institution. Scaling this to commercial property management would need a different control model, and I would rebuild the authorisation layer rather than extend it.
Society administration and revenue operations run on spreadsheets, WhatsApp threads and paper receipts. The product replaces that with a single application covering resident records, billing, collections and notifications.
Production application · product details and links available on request.
Visitor and vehicle access in residential and commercial real estate is a verification problem handled by a person with a register book. The design treats it as a security architecture problem instead.
Labelled as concept and prototype work. It is not deployed, and is presented as architecture rather than as a shipped product.
Four patterns that keep recurring in the work.
Not a stack list. These are the shapes the work takes, and the reasoning behind each one.
How I judge whether an AI project is worth doing at all.
Building is the easy half. The harder judgement is which initiatives justify the investment, which should be piloted, and which should be stopped before they consume a year. Score an initiative below and the framework returns a decision.
An initiative with high value and poor data readiness is not a modelling project — it is a data engineering project with a model at the end. Treating it as the former is the most common way AI programmes lose a year.
A semantic layer or feature store built once serves the next five initiatives. A bespoke model serves one. At portfolio level, reuse potential often outranks the individual business case.
The framework has to be able to say no, or it is decoration. Deterministic rules, a better report, or fixing the underlying process frequently beat a model — and cost a fraction.
Doctoral work, explained the way any other project would be.
A doctorate is a differentiator on an already technical profile, not the headline. What follows is the research framed the way any other case study is framed — problem, method, implication.
Business Administration & AI · Swiss School of Business Management, Geneva · expected October 2026. Stated as candidate until the award is formally conferred.
| Problem | Trust and volatility in decentralised finance — why participants cannot reliably assess counterparty and protocol risk, and what that does to price behaviour. |
| Method | Mixed-method. Exact methodology terminology follows the submitted dissertation. |
| Signals examined | Transparency, governance structures, regulatory posture, and sentiment signals drawn from social sources. |
| Technical connection | The sentiment component sits on the same NLP and retrieval ground as the applied RAG work in this portfolio — the research and the delivery reinforce each other rather than sitting apart. |
| Implication | Intended for platform designers and regulators: which transparency and governance levers actually shift participant trust, as opposed to those that merely signal it. |
Exact title, methodology wording and findings will match the defended dissertation. Findings are not summarised here in advance of the defence.
Thirteen years, described by the work rather than the employer.
Employer names are withheld and available on request. Roles are described by scope, problem and ownership, since that is what transfers.
| Dimension | Evidence |
|---|---|
| People | Teams of up to 10 led directly, with technical direction and delivery accountability. 200+ practitioners trained earlier in career. |
| Technology | Enterprise GenAI analytics platform, fraud and credit decision systems, forecasting pipelines, and a full legacy-to-cloud migration with target architecture. |
| Business | Approximately $1M in estimated fraud losses prevented over three quarters. Pricing decisions moved from negotiation to evidence. |
| Stakeholders | Direct engagement to CXO and director level across sales, pricing, finance, risk and operations. |
| Transformation | Recurring manual analysis converted to automated platforms; a seven-hour nightly legacy estate migrated to orchestrated cloud pipelines. |
Own predictive modelling and analytics automation across pricing and commercial functions, and lead the team delivering it. Built the enterprise GenAI analytics platform. Most of the value comes from turning recurring analytical questions into systems that answer themselves.
Client-facing delivery across financial services, running engagements end to end from problem framing to production handover. Led the fraud engagement that prevented approximately $1M in estimated losses, built the pricing and elasticity work, and executed the migration onto Databricks and Azure. Managed a team of 7–10.
Predictive modelling, forecasting and risk scoring for financial and investment clients, including currency forecasting supporting treasury planning for a Southeast Asian bank. Shipped the interactive applications that put model output in front of decision makers.
Predictive modelling, root cause analysis and ETL automation for banking clients, with the visual analytics layer that made results usable.
Requirements into architecture, data models and delivered applications across finance and media clients. Led a team of four developers.
Mobility applications and dashboards end to end — Python and Java front end, SQL and PL/SQL back end, across banking and insurance data.
Degrees, certifications and the technologies behind the work.
Grouped by what it is used for. Only technologies that can be defended in interview are listed.
What I am looking for, and how to reach me.
Open to Dublin-based Principal AI/ML, AI Architecture, GenAI, Agentic AI and senior AI/Data leadership opportunities. Based in Mumbai, relocating with family for a long-term role.