Back to case studies
SRK Export · Diamond Manufacturing · Colour Prediction AI

How SRK Export Predicts Polished Diamond Colour Before Cutting Begins

How SRK Export Predicts Polished Diamond Colour Before Cutting Begins logo

Built on years of paired rough-to-polish outcomes, the system predicts colour before the stone is cut, and lets the business control how often that prediction runs short.

Controllable limit
Worse-than-predicted rate
Rough / pre-polish
Prediction timing
Matches or surpasses
Accuracy vs. human
Highest from rough
World accuracy claim

At a glance

Overview

Client
SRK Export (Shree Ramkrishna Exports)
Industry
Diamond manufacturing
Domain
Rough-to-polish colour grading and production decision support
Problem
Polished colour is estimated from rough stones by expert eye, and the same stone can get different calls depending on who is looking and when.
Solution
A model trained on multi-year rough-to-polish production history predicts polished colour while the stone is still at rough stage, with the costly "worse than predicted" outcome rate set as a controllable limit rather than a byproduct of judgment.
Scale
Deployed at SRK Export, described in its own materials as the largest diamond manufacturing company.
Result
Colour prediction now happens before cutting and shape decisions are locked in, with the loss-making error rate held lower than manual estimation and tunable to the company's own risk appetite.

Executive summary

World's highest accuracy in polished diamond colour prediction from rough stones: SRK Export's model predicts the polished colour of a rough diamond with the highest accuracy in the world.

Predicts polished colour while the stone is still at rough stage, before cost and shape decisions are locked in.

Matches, and in key cases surpasses, typical human accuracy on the same colour call, validated against historical outcomes.

Keeps the worse-than-predicted ratio low and adjustable, instead of fixed by whoever happens to be grading that day.

Diamond colour is one of the strongest single drivers of a polished stone's value, yet the industry has historically priced that value off a visual estimate made before the stone is even cut. SRK Export built a model that closes that gap: it predicts polished colour from rough-stage and production data, turning years of confirmed rough-to-polish outcomes into a live prediction capability rather than leaving that history as an unused record. The distinguishing design choice isn't the prediction itself; it's what the model is tuned against. Rather than chasing accuracy in the abstract, the system is calibrated specifically to hold down the share of stones that finish worse than predicted, because that is the outcome that actually costs the business money. That ratio is then exposed as a setting the company controls, not a number the model happens to produce.

Project gallery

Project in pictures

01 · Context

Industry Context

Diamond manufacturing runs at real volume on a decision that has no reliable ground truth at the moment it's made: what a rough stone will grade as once it's polished. Colour is one of the primary value drivers in the trade, so this single early call has downstream effects on how a stone is cut, what shape it's planned into, and what price expectation gets attached to it before any of that work is done. The standard approach across the industry has been visual assessment of the rough by experienced graders, and it held for as long as it did for a structural reason: true polished colour can only be confirmed after the stone is cut, so there has never been a way to check the estimate against reality until it's too late to change the plan. What's straining that approach now is scale and continuity: the expertise needed to make these calls well sits with a small number of experienced people, and neither training new staff nor running higher volume closes that gap, because more stones and more people don't make the underlying judgment more consistent.

Manufacturers sitting on years of rough measurements and confirmed polished outcomes have, in most cases, never linked those two records together into something that predicts forward. That's the specific inefficiency a diamond polished colour prediction AI system is built to close: not replacing expert judgment outright, but converting a company's own production history into a repeatable, adjustable prediction layer that sits ahead of the decisions that history used to only explain in hindsight.

02 · Challenge

The Problem

Every working day, manufacturers commit to shape and cost plans based on a colour estimate that two experienced graders can look at the same stone and disagree on. That estimate isn't a stable input; it shifts with fatigue, lighting conditions, and the individual grader's own experience, so the same rough stone can get a different call depending on who evaluates it and when. Because polished colour is only confirmed after cutting, the moment the estimate turns out to be wrong is also the moment the cost and shape decisions built on top of it are already locked in. There's no cheap way to revisit that plan once the mistake surfaces.

The financial exposure here isn't symmetric, which is what makes it structurally dangerous rather than just occasionally inconvenient. If a stone finishes better than predicted, that's an acceptable outcome, even an upside. If it finishes worse than predicted, the stone has effectively been overvalued against the plan built around it, and that's a real loss. Under manual estimation, the share of stones landing in that worse-than-predicted category stays high, and that's the specific ratio that erodes value across a production run, not the average accuracy of the calls.

The industry's reliance on this process is also a scaling problem, not just an accuracy problem. How can AI predict polished diamond colour from rough stones? The answer manufacturers have been sitting on without using it is their own history: paired rough-stage attributes and confirmed polished outcomes, accumulated over years, that could be trained into a predictive model instead of remaining a static archive consulted only after the fact. Most operations generate both halves of that pair, rough data and polished results, but never systematically connect them stone-to-stone at scale.

What made this status quo untenable wasn't a single failure or a compliance deadline; it was the compounding nature of the risk itself. Every stone run through manual-only estimation carries the same asymmetric exposure, and because the worse-than-predicted share stays structurally high under human judgment, that exposure doesn't average itself out over volume; it accumulates with it. That's the pressure that made a systematic, controllable alternative necessary rather than a nice-to-have improvement.

03 · Approach

The Solution

The core decision was to treat this as a supervised learning problem grounded in confirmed outcomes, not an attempt to encode expert reasoning into rules. A rules-based system that mimics how a grader thinks would still be built on the same inconsistent judgment the project was trying to replace. Instead, the model was trained on multi-year historical production data (rough-stage attributes including carat and related rough measurements, process and shape context, and machine and instrument readings already captured on the floor) mapped directly against the actual polished colour grade later observed for that same stone. The prediction target is real outcomes, not simulated expert opinion.

What this makes possible, described as an outcome rather than a feature, is a colour call that arrives at rough stage instead of after cutting, early enough to actually inform the shape and cost decisions it used to only be measured against in hindsight. An off-the-shelf grading tool built for general use couldn't have delivered this, because the entire value of the system comes from being trained on this specific operation's own multi-year rough-to-polish pairing, a dataset that doesn't exist anywhere outside the company that generated it. This had to be built custom because the asset it depends on is proprietary production history, not a general diamond-grading capability that could be licensed off the shelf.

The system was also deliberately scoped to work with fields already captured in the existing production workflow, rather than requiring new sensors, a parallel data-capture process, or a retraining burden on the floor. That's an integration decision as much as a modeling one: the prediction layer sits on top of data that's already being recorded, so adopting it doesn't change how stones move through the plant.

How It Works
1

Rough-stage measurements + process/shape context + instrument readings

2

Model trained on multi-year rough-to-polish outcome pairs

3

Colour grade prediction with confidence score

4

Confidence check

5

High-confidence calls proceed, low-confidence stones route to expert review

6

Colour known before cutting decisions are finalized

04 · Engineering

Technical Deep Dive

01Architecture Brief 01

The genuinely hard part of this project wasn't building a model that predicts a colour grade; it was building one that could be trusted with money on the line, in a domain where the benchmark it's replacing is itself inconsistent. A model can be statistically accurate on average and still be commercially dangerous if it doesn't specifically control for the direction of its errors. Every design decision below traces back to that asymmetry.

  • 01Technical Node

    Rough-to-Polish Outcome Pairing

    The foundation of the system is a dataset that links rough-stage records to their confirmed polished results, stone by stone, across multiple years. This pairing is the actual asset, not the model architecture sitting on top of it: a prediction model trained on unpaired or shallow history couldn't learn the mapping this system depends on. The alternative most operations default to is treating rough data and polished results as separate records that get compared informally rather than linked systematically, which is exactly the gap this project closes.

  • 02Technical Node

    Production-Native Data Ingestion

    The model consumes rough-stage attributes (carat and related rough measurements), process and shape context, and machine and instrument readings that were already being captured on the floor before this project started. No new capture hardware or parallel data process was introduced. The rejected alternative was building a dedicated sensor or intake pipeline for the model, which would have added floor disruption and adoption friction without improving what the model could learn from data that already existed.

  • 03Technical Node

    Supervised Prediction on Confirmed Outcomes

    The model learns the mapping from early-stage signals to actual observed polished colour grades, rather than being built to replicate expert reasoning patterns. Training against confirmed outcomes instead of expert heuristics was the deliberate choice, because expert heuristics are the inconsistent input the project exists to move past; encoding them would have just formalized the same variability the system needed to reduce.

  • 04Technical Node

    Worse-Than-Predicted Risk Calibration

    Rather than optimizing purely for average prediction accuracy, the system is calibrated around the specific rate at which stones finish worse than predicted. This is the single most consequential modeling tradeoff in the build: because an optimistic miss costs the business money and a conservative miss doesn't, tuning for raw accuracy alone would have left that asymmetric risk unmanaged even in a model that looked good on paper. That ratio is exposed as a threshold the company can set, not a fixed output baked into the model.

  • 05Technical Node

    Confidence-Scored Output

    Every prediction returns a colour grade alongside a confidence score, rather than a single unqualified answer. The confidence score is what makes the risk threshold usable in practice. Without it, there's no way to distinguish a call the model is sure about from one it isn't, and no way to route the second kind differently. This prevents the failure mode of treating every prediction as equally reliable.

  • 06Technical Node

    Expert Review Routing

    Low-confidence or borderline predictions are routed to expert review rather than resolved automatically by the model. This keeps human judgment exactly where the model's own uncertainty is highest, instead of fully automating every call regardless of confidence. It's the mechanism that prevents the system from overreaching into decisions it isn't equipped to make alone.

  • 07Technical Node

    Workflow Integration Layer

    The system plugs into stone-field data structures and production systems already in use at the operation, rather than existing as a standalone tool outside normal floor workflow. Fitting into the existing system was treated as a requirement, not an afterthought. A prediction capability that required a new interface or a separate workflow to check would have added friction at exactly the point where speed matters most: before cutting decisions are made.

Tech Stack

Category

Tool / Capability

Why This, Here

Data ingestion

Existing rough-stage measurement and instrument-reading capture

Reuses fields already recorded on the floor; no new capture hardware or parallel process required

Historical foundation

Multi-year paired rough-to-polish outcome records

The stone-level pairing across years is what makes the prediction task learnable in the first place

Model layer

Supervised prediction model trained on confirmed polished-grade outcomes

Learns from actual production results instead of encoding expert heuristics that are themselves inconsistent

Risk calibration

Tunable threshold on the worse-than-predicted ratio

Targets the asymmetric commercial risk directly, rather than optimizing blind average accuracy

Output layer

Colour grade prediction paired with a confidence score

Lets downstream teams treat a high-confidence call differently from a borderline one

Review routing

Expert review queue triggered by low confidence

Preserves human judgment specifically where model uncertainty is highest

Integration

Existing stone-field data structures and production systems

Avoids a standalone tool operating outside the normal floor workflow

05 · Outcomes

Results

Colour prediction now happens at the rough stage, before cutting and shape decisions are locked in, not after.

SRK Export has achieved the world's highest accuracy in predicting polished diamond colour from the rough stone.

Metric

Result

What Changed

Worse-than-predicted ratio

Held lower than the manual-estimation baseline, and set as a controllable limit

The costly outcome type is now a setting the company chooses, not a byproduct of individual judgment

Prediction timing

Rough / pre-polish stage

Colour is assessed before cost and shape decisions are finalized, not after

Accuracy vs. human benchmark

Matches, and in key cases surpasses, typical human prediction

Validated against multi-year historical rough-to-polish outcomes

  • Colour calls move earlier in the pipelinethe estimate now arrives while the stone is still rough, ahead of the shape and cost commitments it used to only be checked against afterward.
  • The costly error type becomes a controlled variablethe worse-than-predicted rate is no longer whatever manual judgment happens to produce; it's a limit the company sets.
  • Expert judgment stays in the loop where it matters mostborderline, low-confidence stones are still routed to human review rather than fully automated.

Beyond the individual figures, what this unlocks is a shift in how the business relates to its own risk. Instead of accepting whatever worse-than-predicted rate manual estimation happens to produce in a given period, SRK Export can now set that rate as a deliberate limit and hold the system to it. The years of rough-to-polish history that used to sit as an unused archive are now doing active work ahead of every cutting decision. And because the confidence score routes uncertain stones back to expert review, the gain doesn't come at the cost of removing human judgment from the process; it comes from applying that judgment only where the model itself signals it's needed.

06 · Process

How We Worked

01Process Step

Establishing the rough-to-polish pairing. Before any modeling work began, rough-stage records had to be linked to their confirmed polished outcomes at the individual stone level across multiple years of production. This pairing, not the model architecture, was the foundation the rest of the project depended on.

02Process Step

Scoping to production-native data. Rather than introducing new capture requirements, the team defined the model's inputs strictly around fields already recorded on the floor: rough measurements, process and shape context, and existing machine and instrument readings. This kept the eventual deployment from requiring any change to how stones move through the plant.

03Process Step

Training against confirmed outcomes. The model was trained to map early-stage signals to actual observed polished colour grades rather than to replicate expert reasoning patterns. This decision protected the system from inheriting the same inconsistency it was built to reduce.

04Process Step

Calibrating for asymmetric risk. Instead of tuning for average accuracy, the team calibrated the model around the worse-than-predicted ratio specifically, and exposed that ratio as a threshold the company could set. This step is what turned a statistically accurate model into a commercially controllable one.

05Process Step

Building the confidence and review path. A confidence score was attached to every prediction, with low-confidence stones routed to expert review rather than auto-decided. This preserved the existing expert layer at the exact point where the model's own certainty was lowest, and made the risk threshold from Step 4 actually enforceable in day-to-day operation.

07 · Insights

Domain Insights

Accuracy is not the metric that protects the business in a domain with asymmetric risk. An optimistic miss (predicting better than the stone actually finishes) costs money, while a conservative miss doesn't. A model can be accurate on average and still be a bad commercial bet if it doesn't separately control the rate of the costly error type; the metric worth calibrating is the worse-than-predicted rate itself, not overall prediction accuracy.

The hard, valuable part of a project like this is the paired history, not the modeling technique. Most operations generate rough data and polished outcomes as separate records and never systematically link them stone-to-stone across years. Building that link is itself the difficult and proprietary part of the work: the prediction model is only ever as good as the pairing behind it, and that pairing doesn't exist by default even in operations that have been recording both sides of the data for years.

The human baseline being replaced is itself an unstable target. Because manual colour calls vary by grader, fatigue, and lighting, there's no single stable "human accuracy" figure to beat; the benchmark shifts depending on who made the call and under what conditions. Any claim that a model "matches or surpasses" human performance in this domain has to be read against that fact: the target itself moves.

08 · Future Scope

Conclusion

SRK Export now gets a polished colour estimate while a stone is still rough, before shape and cost decisions are locked in, and the share of stones that finish worse than that estimate is no longer whatever manual judgment happens to produce on a given day; it's a limit the company sets and holds the system to. That's the transformation in plain terms: a subjective, hindsight-only estimate became a forward-looking, adjustable one, built entirely out of the company's own production history. This is what delivers the world's highest accuracy in polished colour prediction from rough diamonds.

The same paired dataset and prediction approach extends naturally to other polished attributes beyond colour, and the risk-calibration pattern, tuning to a controllable costly-error rate rather than raw accuracy, applies to any other stone-grading decision in the pipeline where an optimistic miss and a conservative miss don't cost the same. The years of history a manufacturer already has are worth more as a live prediction layer than as an archive.

Frequently Asked Questions

By training a supervised model on multi-year paired records that link rough-stage attributes (carat and related measurements, process and shape context, and instrument readings) to the confirmed polished colour grade later observed for the same stone. The prediction target is real production outcomes, not encoded expert opinion, so the model learns the mapping manufacturers already generate but rarely connect stone-to-stone at scale.

It is tuning the system around the rate at which stones finish worse than the colour call, not around average accuracy alone. An optimistic miss costs money; a conservative miss does not. SRK Export exposes that ratio as a threshold the company can set, so the costly error type becomes a controllable limit rather than a byproduct of whoever graded the stone that day.

A rules-based system that mimics how a grader thinks would still be built on the same inconsistent judgment the project was trying to replace. Expert heuristics vary by fatigue, lighting, and individual experience. Training against confirmed polished outcomes avoids formalizing that variability and learns from results the business already knows are true.

No. Every prediction includes a confidence score, and low-confidence or borderline stones are routed to expert review rather than auto-decided. Human judgment stays exactly where the model's uncertainty is highest, while high-confidence calls proceed early enough to inform cutting and shape decisions.

The value comes from training on SRK Export's own multi-year rough-to-polish pairing, a proprietary production history that does not exist outside the company that generated it. A general grading product cannot deliver the same mapping or the tunable worse-than-predicted threshold, because those depend on this operation's confirmed outcomes and existing floor data.

Want similar results in your production line?

Share your constraints and targets. We'll propose an automation roadmap with measurable quality and throughput outcomes.