How DataBounty works

DataBounty is a marketplace for coding datasets. Sponsors request and fund them, builders create the items, and validators audit them. An automated verification pipeline sits in the middle so that only work that provably matches the spec gets accepted and rewarded.

Coding is the live domain today; legal, healthcare, finance, and math/science are opening next. see_all_domains →

Live today: open datasets, paid in karma

Claim batches with no bond, earn karma per accepted item, and get named credit when eligible work publishes to Hugging Face. Paid bounties, funded in USDC and paid per accepted item, are coming soon. Karma holders get first claim. how_karma_works →

01

The loop

From spec to delivered dataset in ten steps. Every step is funded, verified, and auditable.

01 create_bounty02 llm_plan03 fund_pilot_tranche04 pilot_review05 fund_full_tranche06 creators_claim_batches07 validation_pipeline08 independent_audits09 accepted_work_paid10 dataset_delivered
  1. 01
    Create the bounty

    The sponsor specs the dataset: category, language, item count, difficulty mix, license type, and deadline. Nothing is public yet.

  2. 02
    LLM plan & platform review

    The platform drafts a structured plan (slots, per-item rewards, pilot size) and reviews the spec for feasibility before it can be funded.

  3. 03
    Fund the pilot tranche

    Roughly 1-2% of the total budget is deposited in USDC. This opens a small pilot batch to trusted contributors, so exposure stays tiny.

  4. 04
    Pilot review (7-day auto-approve)

    The sponsor inspects real accepted pilot items. They can approve, request one spec revision, or walk away. Silence auto-approves after 7 days.

  5. 05
    Fund the full tranche

    The spec locks and the remaining budget moves to escrow. From here, spec-matching work cannot be rejected on taste.

  6. 06
    Contributors claim batches

    Contributors claim 10-item batches inside slots, each with a clear reward per item and a 72-hour deadline. No bond while on time.

  7. 07
    Validation pipeline

    Every submission runs the automated gauntlet: duplicate wall, benchmark-contamination screen, sandboxed test execution, LLM validation.

  8. 08
    Independent audits

    Depending on audit mode, ranked validators review some or all provisionally accepted items and flag real issues for bonuses.

  9. 09
    Accepted work is paid

    Accepted items trigger USDC payouts from escrow to contributor wallets; validators earn base rewards plus confirmed-issue bonuses.

  10. 10
    Dataset delivered

    The export is handed to the sponsor under the chosen license, with quality stats and a public sample published on the marketplace.

02

The verification model

Five layers between a submission and a payout. An item must clear all of them, with no exceptions and no manual overrides.

layer 1

Duplicate wall

Each submission is similarity-scored against everything already accepted in the bounty, and across the marketplace. Near-duplicates are rejected before they cost anyone review time, and they are never paid.

layer 2

Contamination screening

Items are checked against public coding benchmarks and well-known problem sets. A dataset that leaks eval questions is worthless for training, so likely benchmark copies are flagged and blocked.

layer 3

Sandboxed test execution

For execution-verified types (debugging, implementation, SQL, regex, translation, performance) the contract is executable: broken code must fail the submitted tests and fixed code must pass all of them. If the checks do not hold in the sandbox, the item bounces automatically.

layer 4

LLM validation

A model reviews each item against the bounty spec (prompt clarity, solution correctness, difficulty calibration, explanation quality) and produces a pass score with notes.

layer 5

Human audit (per audit mode)

Bounties choose LLM-only, partial (25% sample), or full audit. Ranked validators re-review provisionally accepted items and earn a bonus for every confirmed issue, a direct incentive to find real problems.

AI-assisted creation is allowed

We verify the output rather than the process. Contributors may use AI tools, or generate items entirely with AI, as long as the generation method is disclosed on each submission. Every item faces the same pipeline either way, and each bounty publishes its human / AI-assisted / AI-generated mix so buyers know exactly what they are getting.

03

Pilot & tranche funding

Nobody commits the full budget to an untested spec. The pilot is the cheapest possible way to find out if the spec is right.

~1%

Pilot tranche first

Only about 1-2% of the budget is deposited to open the pilot. A small batch is produced by trusted contributors and runs the full pipeline, so the sponsor judges real, verified samples.

7 days

Review with auto-approve

The sponsor has 7 days to approve the pilot or request one spec revision. If they do nothing, the pilot auto-approves, so contributors are never left waiting on an absent buyer.

Spec locks after pilot

Once the full tranche is funded, the spec is frozen. Work that matches the locked spec and passes the pipeline cannot be rejected. Taste disputes belong in the pilot phase, before contributors have done the work.

04

Money & fees

Everything settles in USDC on Solana. The fee schedule and escrow rules are the same for every bounty.

Platform fee

20%
standard · 30-day hold
25%
extended · 12-month hold

The fee covers the validation pipeline, sandbox execution, escrow operations, and dispute resolution. Every delivered dataset can later be listed in the resale catalog; the license only sets how long it is held off the market first: 30 days on Standard, 12 months on Extended. Contributors keep a 50% pro-rata resale royalty. The fee is carved out of the budget up front and shown on every bounty page.

Escrow in the platform treasury

Deposits sit in the platform treasury, reserved per bounty into contributor pool, validator pool, and fee. Rewards accrue as items are accepted and are paid out in periodic USDC batches to contributor and validator wallets.

Partial-fill policy

If a bounty ends below target, every accepted item is still paid in full. The sponsor takes delivery of the partial dataset at a prorated price and the unspent escrow is refunded. Underfilling is a priced outcome, not a dispute.

Bonds only after abandonment

There is no upfront stake. A contributor or validator who abandons a claimed batch must post a bond (10% of the expected reward) on future claims until their record recovers. Deliver on time and you never lock a token.

05

Ranks

Ranks are earned per role. Higher ranks unlock capacity and access, but they are never a substitute for passing the pipeline.

</> builder_ranks (contributors)
1 Scout1 active batch
2 Apprentice2 active batches
3 Builder30-item batches
4 Specialist75-item capacity
5 Craftsmanpriority batch access
6 Senior Buildersenior capacity
7 Expertexclusive bounties
8 Architectarchitect capacity
9 Principalspec consultation invites
10 Master Buildertop payout priority
◇ auditor_ranks (validators)
1 Observersmall audit batches
2 Reviewerstandard batches
3 Inspectorstandard batches
4 Auditorlarger + bonus mult.
5 Senior Auditorfull-audit bounties
6 Quality Leadquality-lead access
7 Verifierverifier duties
8 Principal Verifierdispute-panel access
9 Arbiterarbitration duties
10 Master Arbitertop bonus multiplier
06

Common questions

The short version of the rules everyone plays by.

Who pays for a bounty?

The sponsor: a lab, company, or individual who wants the dataset. They deposit USDC on Solana into the platform escrow in two tranches, a small pilot first, then the full budget after they approve real pilot samples.

Who earns, and how much?

Contributors earn a fixed USDC reward per accepted item (set per slot). Validators earn a base reward per audit batch plus a bonus for each confirmed issue they flag. All rates are published on the bounty before anyone claims work.

What happens if a bounty underfills?

Partial fill is a first-class outcome. All accepted items are paid at the full per-item rate, the sponsor receives the partial dataset at a prorated price, and the unspent escrow is refunded. Contributors never eat the shortfall.

Is AI-generated work allowed?

Yes. AI assistance and full AI generation are allowed, with disclosure. We verify the output rather than the process: every item faces the same duplicate, contamination, execution, and validation checks regardless of how it was made. Each bounty publishes its generation mix.

Can the sponsor reject work they simply don't like?

No. After the pilot is approved, the spec locks. Items that match the locked spec and pass the pipeline are accepted and paid from escrow. The pilot phase exists precisely so taste disagreements get resolved before the full budget is committed.

When is KYC required?

Browsing, claiming, and small earnings are wallet-only. Once cumulative payouts to a wallet cross the regulatory threshold (currently $600), payouts pause until identity verification is completed. This is shown in your dashboard well before it applies.

Ready to see it in motion?

Browse the open datasets or head to the dashboard to start.