VIDA Digital Identity logo

Fraud / Device Intelligence Data Scientist

VIDA Digital Identity
1 day ago
Full-time
On-site
East Jakarta, Java, Indonesia
Data Analyst

About the Role

VIDA verifies identity in real time for financial services, lending, telecommunications, and government-linked services. This makes us a target for organised fraud operations that spoof genuine users and devices at scale — and adapt within days when a defence starts working. Detection here is a continuing contest, and errors cost in both directions: missed fraud harms users and customers, while false positives deny real people access to services.

You will own detection on this path: investigate attacks in raw device and behavioural telemetry, build the rules and ML models that catch them, and measure what they catch and what they cost genuine users. The data is high-dimensional and noisy, labels are scarce and delayed, and decisions must be explainable.


What You Will Do


Understand the adversary

  • Investigate real attacks — reconstruct how a campaign operates from the evidence available, and characterise it well enough that we can act on it.
  • Find signal in high-dimensional, messy data — work across large volumes of device and behavioural telemetry to identify what genuinely separates fraudulent activity from legitimate users.
  • Keep pace with change — when an adversary adapts, quickly establish what changed and what still distinguishes them.


Build detection that holds up

  • Design detection logic and models — from feature engineering through training, calibration, and threshold selection, choosing the right tool for the problem: interpretable logic where explainability matters, models where they earn their place.
  • Detect coordinated activity, not just single events — move detection beyond judging each request in isolation toward recognising organised behaviour across many, which is far harder for an adversary to disguise.
  • Balance detection against user impact — treat the false-positive cost to genuine users as a first-class constraint, measured on the real population before anything is enforced.
  • Own quality in production — validate before enforcement, monitor for drift and degradation, and know when to roll back.


Make the evidence trustworthy

  • Build and improve ground truth — establish reliable labels in a domain where they are scarce, delayed, and imperfect, and be clear about what your labels do and do not cover.
  • Hold a high evidential bar — recall and false-positive rates measured on the right populations, class imbalance, selection bias, and correctness of point-in-time evaluation. You should be able to spot the ways an analysis can flatter itself, and design around them.
  • Make the work reproducible and legible — analysis and specifications that colleagues can re-run, challenge, and build on.
  • Partner closely with fraud operations, engineering, and product — and communicate findings clearly to non-specialists, including when the honest answer is uncomfortable.


What We Are Looking For


Must-have

  • 4+ years of applied data science or machine learning in fraud, risk, abuse, trust & safety, security, or a closely adjacent adversarial domain — with work you took to production and measured.
  • Strong SQL on large datasets — comfortable and effective in large, messy event data.
  • Fluent Python data-science stack — pandas, scikit-learn, and gradient-boosted trees, with clean and reproducible analysis.
  • Genuine evaluation rigour — you understand why a model can look excellent offline and fail in production: label leakage, selection bias, drift, class imbalance, and the difference between conditioned and population-level error rates.
  • Feature engineering from raw, semi-structured data — turning noisy, high-cardinality inputs into signals that hold up over time.
  • An adversarial mindset — you instinctively ask how someone would evade what you built, and you design accordingly.
  • Sound judgement on explainability — decisions affecting a person's access to services need to be defensible, not just accurate.
  • Ownership and pragmatism — you can carry a problem from ambiguous question to shipped, measured outcome without waiting to be handed a specification.


Nice-to-have

  • Device intelligence, device fingerprinting, or bot and automated-abuse detection.
  • Biometrics, liveness, or presentation- and injection-attack detection.
  • Mobile platform knowledge, including device integrity and attestation concepts.
  • Graph analysis or entity resolution for linking related activity and identifying organised rings.
  • Real-time or streaming features and the discipline of keeping training and serving consistent.
  • Experience in emerging markets or high-growth financial services.
  • Comfort using AI-assisted tooling to work faster without lowering the evidential bar.



What are we trying to solve?

We have 7.5 billion people on Earth, of which over 1 billion cannot securely prove their identity right now. Every year, 140 million babies are born, of which 40 million go unregistered. Simply put, these people are deprived of social benefits, such as education and health, their civil rights to vote and travel; and are excluded from the economy because they cannot sign up for bank accounts, loans, welfare programs etc. We believe this is unacceptable, and needs to change.

At VIDA, we are creating a frictionless digital identity system. One that fulfils the needs and expectations of our times, and is available anywhere, for everyone.


Why are we solving this problem?

The United Nations (UN) and World Bank ID4D initiatives aim to provide everyone on the planet with a legal identity by 2030. This deadline is just a few years away, we are expecting a digital identity to be a legal human right by then and we at VIDA want to be pioneers in leading this change.


Who are we?

We are a highly driven bunch of people solving this problem for our own reasons. Whether it is misleading doctors, or because we didn't get access to fair ration due to corruption — our collective goal aligns.