Back to case studies
SRK Exports · Diamond Manufacturing · Clarity Grading AI

Automating Diamond Clarity Grading at SRK Exports

Automating Diamond Clarity Grading at SRK Exports logo

A multi modal grading system that respects ordinal grade relationships, matches human-level accuracy, keeps training itself automatically, and runs live inside SRK's production quality control workflow.

Human-level
Accuracy level
Auto-training
Improvement path
1 TB+
Production data
30+
Cut shapes

At a glance

Overview

Client
SRK Exports (Shree Ramkrishna Exports Pvt. Ltd.)
Industry
Diamond manufacturing and export
Domain
Automated diamond clarity grading for production quality control
Problem
Manual clarity grading caps how many stones a facility can process per shift and produces inconsistent calls at borderline grade boundaries.
Solution
An ordinal, multi modal system that predicts clarity grade and confidence from video, images, and inclusion data, live inside the QC workflow.
Scale
1 TB+ of real production data, 30+ cut shapes
Result
A live production system returning grade and confidence scores in real time, currently at human-level accuracy and still improving through automatic training.

Executive summary

A single grade difference on a polished diamond, VS1 instead of VS2, can move its price by thousands of dollars, yet the judgment behind that grade has stayed almost entirely manual for decades. SRK Exports, a diamond manufacturer that has cut, polished, and graded stones in house since 1964, built a diamond clarity grading AI system to support its own production quality control, a workflow that runs across more than 2,000 master artisans and its own internal grading standard. The system respects the ordinal structure of clarity grades, fuses video, still images, and structured inclusion data instead of relying on any single input, and now runs live in production, returning a predicted grade and confidence score in real time. Even graders with 30+ years of experience find this task difficult, and the model has now reached human-level accuracy on it. It is still training automatically, and its accuracy is expected to keep increasing.

Project gallery

Project in pictures

01 · Context

Industry Context

The diamond industry grades millions of polished stones every year, and at every stage, internal manufacturing QC, independent laboratory certification, and retail, each stone gets examined under 10x magnification, its inclusions identified and mapped, and a clarity grade assigned from an ordinal scale of 11 official grades. That grade is not a cosmetic label. It directly sets price, so a single step between VS1 and VS2, or VS2 and SI1, can shift a stone's value by thousands of dollars. Because of that, the industry has spent years looking for a diamond clarity grading AI system that can operate at the volume large-scale production actually demands, without losing the judgment that makes a grade trustworthy.

The manual approach held for as long as it did because clarity grading genuinely requires expert judgment: a grader has to weigh size, number, position, nature, and relief together, across video rotation, still captures, and physical attributes, and arrive at one ordinal call. Even a grader with 30+ years of experience does not find this easy, and building that level of judgment takes decades of practice. No off-the-shelf computer vision or tabular model was built to do that. What is breaking the manual only model now is volume. A large manufacturer grading thousands of stones per day runs into a hard ceiling: trained graders are expensive to develop, difficult to scale quickly, and subject to fatigue and inconsistency in ways that cap how fast a production line can move.

02 · Challenge

The Problem

Every working day, a large manufacturing facility pushes thousands of polished stones through clarity inspection. Each stone needs a trained grader to examine it under magnification, map its inclusions against five formal criteria (size, number, position, nature, and relief), and commit to one of 11 ordinal grades. The hardest calls cluster at specific boundaries, VS2 versus SI1, SI1 versus SI2, where even experienced graders frequently disagree, and those borderline stones make up a meaningful share of total production volume, not an edge case.

The cost of getting it wrong is not abstract. A stone graded VS2 instead of SI1 can be undersold by 10 to 20 percent of its value. A stone graded SI1 instead of VS2 creates a trust problem the moment a buyer notices the inclusion later. Inconsistent grading across batches compounds both issues: it erodes buyer confidence and makes inventory harder to manage at scale. And because manual grading throughput sets a hard limit on how many stones a facility can move per shift, the bottleneck is not just a quality issue, it is a production capacity issue.

That points directly at the operational question manufacturers have been asking for years: how can diamond manufacturers automate clarity grading at scale without losing the judgment a trained grader brings to a borderline stone? The honest answer, until recently, was that they could not. The bar is set by people with 30+ years of experience, and even at that level, consistent accuracy is hard to achieve. Clarity grading is a multi dimensional, ordinal, expert judgment problem, and generic image classifiers cannot capture ordinal grade relationships, tabular models miss the visual evidence, and single-photo approaches ignore how inclusions appear differently as a stone rotates under a loupe.

There was no single dramatic failure that forced the issue. What accumulated instead was structural pressure: production volume at scale manufacturers had already outrun what manual grading could sustainably support, and no existing method was built to handle the actual shape of the problem, ordinal grades, multi modal evidence, more than 30 shape variations, and rare grade imbalance, all at once. The industry wanted automated, consistent, scalable clarity evaluation for a long time. What kept it out of reach was the absence of a large enough body of real graded data, and a system built for the problem's real complexity rather than a simplified version of it.

03 · Approach

The Solution

The starting decision was to treat clarity grading as an ordinal problem rather than a standard classification task. A model that scores all mistakes equally would treat a VS1 stone predicted as VS2 the same way it treats a VS1 stone predicted as SI1, even though the second error is far more commercially damaging. The team rejected standard multi class classification for that reason and built a model that penalizes distant grade errors more heavily than adjacent ones, matching how the industry actually experiences a misgrade.

The second decision followed from the first: rather than forcing a single hard label on every stone, the system outputs a probability distribution across clarity grades, along with a directional signal for borderline stones indicating whether a stone leans toward the grade above or below. Forcing a binary decision on a genuinely ambiguous VS2/SI1 stone would just relocate the disagreement problem graders already have, not resolve it. Giving graders a probability and a direction instead lets them apply their own judgment exactly where it is most needed.

What is now possible is a production floor where every stone gets a consistent, data-driven starting assessment in real time, at human-level accuracy, before a human grader ever commits to a final call. Building that required custom work at almost every layer: fusing video, still images, structured inclusion annotations, and physical attributes (carat, shape, color) into one model, since a system built on images alone misses structured grading context, and a system built on tabular scores alone misses the visual evidence a grader actually relies on. An off-the-shelf vision model was never going to handle that fusion, the ordinal grade structure, and 30+ shape variation at once.

The system was designed to fit inside the existing QC workflow rather than replace it. A stone still gets a full inspection video and grader-marked inclusions the same way it always did; the model simply processes that same inspection package and returns a grade and confidence score before the grader finalizes their assessment. The model also does not stand still: it trains automatically as new graded stones move through production, so its accuracy is set to keep increasing over time.

How It Works
1

Stone arrives at QC inspection

2

Inspection video and images captured, inclusions marked by grader

3

Structured grading data recorded (carat, shape, color, zone scores)

4

Model processes video, images, annotations, and attributes together

5

Predicted clarity grade and confidence score returned

6

Grader reviews the prediction and finalizes the grade

7

Stone proceeds through the production pipeline

8

Finalized grades feed back into automatic model training

04 · Engineering

What Makes This a Big Deal

01Architecture Brief 01

This project is not big because of any single component. It is big because several hard things came together in one working system: a task that takes humans 30+ years to master, a model that has already reached that level, a system that keeps improving without manual retraining cycles, and a deployment that runs live on a real production floor.

  • 01Technical Node

    Human-Level Accuracy on a Task That Takes Decades to Master

    Clarity grading is one of the hardest judgment calls in the diamond industry, and the model has reached human-level accuracy on it. The graders whose work sets the benchmark have 30+ years of experience, and even at that level, accuracy is difficult to keep consistent because the calls at borderline boundaries are genuinely subjective. Getting a model to that level is not routine: most automated approaches stall well short of it. Reaching it is the core milestone of this project.

  • 02Technical Node

    A Model That Trains Itself and Keeps Improving

    Human-level accuracy is where the model is today, not where it stops. The model is set up for automatic training, so it keeps learning as new graded stones pass through the production workflow. Accuracy is expected to keep increasing over time, without a rebuild from scratch and without waiting for manual retraining cycles. A human grader's skill takes decades to build and stays with one person, while this system's accumulated skill compounds across every stone it sees.

  • 03Technical Node

    Built on 1 TB+ of Real Production Data

    The system learns from real, in-house inspection data, not synthetic or publicly sourced data. More than 1 TB of production video, images, and inclusion records were unified into a single training and inference pipeline. That volume is what allows the model to learn across the full commercial clarity range, including the rare high grades that a smaller dataset would barely touch.

  • 04Technical Node

    One System for 30+ Cut Shapes

    A single deployed system serves the full shape mix a production floor actually manufactures. Round brilliant, emerald, pear, cushion, marquise, princess, and other fancy cuts are all covered, rather than requiring a separate tool for each shape. For a manufacturer, that means one grading system across the whole line, not a patchwork of tools.

  • 05Technical Node

    Live in Production, Working in Real Time

    This is not a prototype or a research demo. The system runs as a live service inside SRK's manufacturing QC workflow, returning a predicted grade and confidence score in real time while a stone is still moving through inspection. It was built to work alongside graders in line, so it adds capacity to the workflow instead of creating a new bottleneck.

  • 06Technical Node

    Consistency and Capacity at Scale

    Every stone now gets the same standard of assessment, every time. Human graders are subject to fatigue and batch-to-batch inconsistency; a model applies one consistent standard from the first stone of the day to the last. That reduces the burden on graders for the large volume of straightforward stones and frees their attention for borderline and exceptional cases, which directly raises how many stones a facility can move per shift.

At a Glance

Category

Capability

Why It Matters

Accuracy

Human-level, benchmarked against graders with 30+ years of experience

The model can be trusted as a starting assessment on a task that is difficult even for experts

Learning

Automatic training, continuously improving

Accuracy keeps rising without manual retraining effort

Data

1 TB+ of real in-house production data

Grounded in the actual stones and conditions the system operates on

Coverage

30+ cut shapes in one system

Serves the full production mix without separate tools per shape

Output

Grade, probability, and directional borderline signal

Shows real ambiguity instead of forcing false confidence

Integration

Live production service inside the QC workflow

Returns grade and confidence in real time without disrupting inspection

05 · Outcomes

Results

The model has reached human-level accuracy on diamond clarity grading, a task where graders with 30+ years of experience still find consistent accuracy difficult. It is live in production, training automatically, and still improving.

Metric

Result

What It Meant Operationally

Accuracy level

Human-level accuracy reached

The model's grading matches the level of experts with 30+ years of experience

Improvement path

Automatic training, still improving

Accuracy is expected to keep increasing as the model continues to learn from new production stones

Total data volume

1 TB+

Forced a streaming, cache-based pipeline rather than a simpler in-memory approach

Shape coverage

30+ cut shapes

One deployed system now serves the full shape mix a production floor actually manufactures

Deployment status

Live in production

Runs inside the manufacturing QC workflow today, not as a prototype

  • Human-level accuracy already reached on a task that takes even 30+ years of experience to perform well, and where consistent accuracy is difficult for people as well.
  • Automatic training, so accuracy keeps increasing the model is not a one-time build; it keeps learning and getting better.
  • 1 TB+ of production data unified into a single training and inference pipeline this is real, in-house inspection data, not synthetic or publicly sourced.
  • 30+ cut shapes handled by one system round brilliant, emerald, pear, cushion, marquise, princess, and other fancy cuts are all covered rather than requiring separate tools per shape.

Beyond the numbers, what this unlocked is a QC workflow where every stone now gets a consistent, data-driven starting assessment in real time, at human-level accuracy, before a grader commits to a final call. That reduces the burden on graders for the large volume of straightforward stones and lets their attention concentrate on the borderline and exceptional cases that genuinely need human judgment. And because the model continues to train automatically, the gap between what it can do today and what it will do tomorrow is expected to keep widening in its favor.

06 · Process

How We Worked

01Process Step

Assembling a production dataset. The team pulled together 1 TB+ of individually graded stone data directly from SRK's live grading pipeline: inspection video, still images, inclusion annotations, and physical attributes. Using real in-house data, rather than synthetic or public sources, meant the system learned from the same conditions it would later operate in.

02Process Step

Designing the ordinal, multi modal model. The team built a model architecture that both respects the ordinal structure of the 11 clarity grades and fuses video, still images, inclusion annotations, and physical attributes into one prediction. This was the core decision the rest of the system depended on.

03Process Step

Building the directional borderline output. Rather than stopping at a single predicted grade, the team added a probability distribution and directional signal specifically for borderline stones near the VS2/SI1 and SI1/SI2 boundaries. This gave graders a genuinely useful signal on exactly the stones where human disagreement is highest.

04Process Step

Deploying into the live QC workflow. The finished system was integrated directly into SRK's manufacturing quality control process, returning a predicted grade and confidence score in real time as stones move through inspection. This step turned a trained model into a production tool that graders actually use, without changing how a stone gets inspected in the first place.

05Process Step

Enabling automatic training. With the system live, the model was set up to keep training automatically as new graded stones flow through production. This is why human-level accuracy is a starting point rather than a ceiling: the model continues to improve on its own.

07 · Insights

Domain Insights

Clarity grading is fundamentally an ordinal problem, not a classification problem, and that distinction matters more than it might first appear. Treating a VS1-to-VS2 miss the same as a VS1-to-SI1 miss misrepresents how this industry actually experiences error: one is commercially trivial, the other can cost real money and buyer trust. Anyone approaching a grading problem like this needs to design the loss function and evaluation metric around that ordinal structure from the start, not bolt it on afterward.

No single evidence type is sufficient on its own in this domain. Video, still images, structured annotations, and physical attributes each carry information the others miss, and a system that leans on just one of them will systematically miss what a human grader would have caught. Multi modal fusion here is not an enhancement, it is a requirement for matching how a grader actually works.

Human-level is a genuinely hard bar in this domain. Graders with 30+ years of experience have built their judgment through decades of looking at stones, and even they do not always agree on borderline calls. Any automated system should be measured against that standard, not against a simplified benchmark, and it should be built to keep learning, because expert-level judgment is never finished. That is why automatic training matters as much as the first version of the model.

08 · Future Scope

Conclusion

SRK Exports now runs a live, ordinal, multi modal system inside its own manufacturing QC workflow, returning a grade and confidence score in real time on every stone that moves through inspection. The model has reached human-level accuracy on a task where even graders with 30+ years of experience find consistent accuracy difficult, and it continues to train automatically, so its accuracy is expected to keep increasing. What changed is not that human graders were replaced, it's that every stone now gets a consistent, data-driven starting point before a grader makes the final call, and graders can spend their attention on the borderline and exceptional stones that actually need it.

The same fused evidence pipeline, video, images, annotations, and physical attributes, now exists as a foundation that keeps growing stronger with every stone it sees, and that can extend further into rare high-clarity grades as more data accumulates, without a rebuild from scratch. Grading a diamond has always come down to one thing: knowing exactly which imperfections matter, and by how much.

Frequently Asked Questions

By treating clarity as an ordinal, multi modal problem rather than a standard image classification task. SRK Exports fused inspection video, still images, inclusion annotations, and physical attributes into one model that returns a grade, probability distribution, and directional borderline signal, then deployed it live inside the existing QC workflow so graders still finalize the call.

Because not all mistakes cost the same. Predicting VS2 for a VS1 stone is commercially different from predicting SI1 for that same stone. An ordinal model penalizes distant grade errors more heavily than adjacent ones, matching how the industry experiences a misgrade, instead of scoring every wrong label equally like a standard multi-class classifier.

No. The system returns a predicted grade and confidence score as a consistent starting assessment while the stone is still in inspection. Graders review the prediction, especially on borderline stones where the model provides a directional signal, and finalize the grade. The goal is capacity and consistency, not removing expert judgment from the hardest calls.

Finalized grades from production feed back into automatic training as new stones move through the QC workflow. Human-level accuracy is the current starting point, not a ceiling: the model keeps learning without waiting for manual retraining cycles or a rebuild from scratch.

A vision-only model misses structured grading context; a tabular-only model misses the visual evidence a grader actually relies on. Clarity grading also needs ordinal loss design, 30+ cut-shape coverage, and rare-grade imbalance handling. That fusion had to be purpose-built on 1 TB+ of real in-house production data.

Want similar results in your production line?

Share your constraints and targets. We'll propose an automation roadmap with measurable quality and throughput outcomes.