Service · Remote-first senior specialists

Blockchain & On-Chain Data Analytics

Raw chain data is not analytics. MV.tech decodes transactions, logs and events using contract ABIs and Web3.py, consolidates them across networks into BigQuery, and adds the valuation, categorisation and reconciliation layers that let finance, product and compliance teams answer questions about on-chain activity.

  • Remote from Ahmedabad, India
  • Overlap with your timezone
  • Senior specialists, direct access
  • Scope agreed in writing

The starting point

When on-chain activity has to answer business questions

  • A block explorer can show one transaction, but nobody can total a quarter across chains.
  • Each network has its own ingestion script and its own idea of what a transaction record looks like.
  • Finance needs token movements valued in fiat at the right timestamp, with fees attributed.
  • Staking rewards are reconciled by hand and the totals never quite agree.
  • A compliance or AML question needs wallet-level history and there is no dataset to query.
  • Someone corrected a figure directly in the analytics table and the original value is gone.
  • Token coverage is partial, so a material share of transfers lands in an "unknown" bucket.

What we deliver

On-chain data engineering services

01

Cross-chain ingestion and ETL

Pull blocks, transactions, logs and events from node RPCs, indexers and chain APIs into a common schema in BigQuery. Reorg handling, backfills, checkpointing and per-chain rate limits are part of the pipeline rather than an afterthought. Work has covered Ethereum, Bitcoin, Immutable X, Immutable zkEVM, XDC, WEMIX, Tezos and Cardano.

02

Decoding transactions, logs and events

Apply contract ABIs with Web3.py to turn opaque input data and topics into named events with typed arguments: transfers, mints, burns, stakes, swaps. That includes proxy contracts, non-standard implementations and the contracts whose interface has to be reconstructed before anything can be read.

03

Token identification and valuation

Maintain token registries and price reference data, and value movements at the relevant block timestamp in the reporting currency, handling decimals, fee proration and FX explicitly. Widening mapping coverage is ordinary, high-value work: on past engagements it moved from 35 tokens to 42 and beyond, pulling material volume out of the unknown bucket.

04

Transaction categorisation

Classify activity into the categories the business reports on — transfer, trade, fee, reward, bridge, internal movement — using contract, method, counterparty and pattern rules that are written down and reviewable rather than buried inside a query.

05

Staking-reward reconciliation

Reward ledgers per validator, delegator or wallet, reconciled against on-chain distributions and internal records, with period totals, accrual timing and a discrepancy report. On past engagements a ledger of this kind has removed more than 70% of a manual reconciliation cycle.

06

Wallet analytics and attribution

Wallet-level histories, balances over time, cohort and retention views, and attribution of addresses to known entities or internal accounts, with the labelling maintained as data rather than hardcoded into reports.

07

Compliance and AML investigative datasets

Queryable datasets for investigation and monitoring: counterparty exposure, flows through addresses of interest, thresholds and time-windowed aggregates. They support an analyst’s judgement and an auditor’s questions; they do not replace either.

08

Immutable base data with an audit trail

Base tables hold exactly what the chain reported and are never edited. Manual corrections, whitelists and overrides live in separate layers applied when the reporting view is built, each carrying an author, a timestamp and a reason, so a figure can still be explained months later.

  • Python
  • Web3.py
  • Contract ABIs
  • BigQuery
  • Node RPCs
  • Ethereum
  • Bitcoin
  • Immutable X
  • Immutable zkEVM
  • XDC
  • WEMIX
  • Tezos
  • Cardano
  • Pandas
  • Parquet
  • SQL
  • Docker
  • Looker Studio

A useful first project

A useful first project: one chain, end to end

Take a single network and one business question — a quarterly token flow, a staking-reward total, a wallet cohort — and build it the whole way: ingestion, decoding with the relevant ABIs, token mapping, valuation, categorisation and a reconciled output compared against whatever figures exist today. One chain done properly establishes the schema, the decoding patterns, the override model and the validation checks, so the second and third networks are mostly configuration rather than new engineering.

See our delivery process

Working together

A practical path from scope to delivery.

  1. 01

    Agree the question and the chains

    Identify the networks, contracts and wallets in scope, the reporting currency and period, and the figures the output has to reconcile against.

  2. 02

    Ingest, decode and value

    Build the pipeline and decoding layer, map tokens, apply valuation and categorisation, and check coverage and gaps block by block.

  3. 03

    Reconcile and operate

    Compare against on-chain and internal sources, document the override path, then schedule, monitor and hand over with the mapping rules.

Teams our engineers have worked with

  • Google
  • Volvo
  • BCW
  • RootstockLabs
  • Chainlabs
  • Toptal
  • Turing

Before we begin

Questions about blockchain & on-chain data analytics.

Not answered here? Ask us directly or read the full FAQ.

Why not just use a public chain dataset?

Public datasets are a reasonable starting point for common chains and standard token transfers. They tend to fall short on newer or smaller networks, on contracts whose events are not standard, and on the business layer: your token registry, your categorisation rules, your reporting currency and your reconciliation against internal records. We use public datasets where they fit and build ingestion where they do not.

How do you handle a contract with no published ABI?

Several ways, in order of preference: a verified ABI from the explorer, a standard interface the contract conforms to, an ABI supplied by whoever deployed it, or reconstruction from method signatures and observed event topics. Whichever route is used, the decoding rules are recorded per contract and the raw input data is retained, so decoding can be corrected and replayed without re-ingesting the chain.

Can analysts correct a figure without breaking the audit trail?

That is what the override layer is for. Base tables hold what the chain reported and are never edited. Corrections, whitelists and manual classifications live in separate tables with an author, a timestamp and a reason, and are applied when the reporting view is built. Both the original and the adjusted figure stay available, which is usually what an auditor asks to see.

How do you value tokens that barely trade?

Carefully, and visibly. The valuation method is a documented rule per token: reference price at block timestamp, a defined source hierarchy, a fallback, or explicitly no valuation at all. Unvalued volume is reported rather than hidden. It is better for a finance team to see that a share of activity has no reliable price than to receive a confident number built on a thin market.

Can this feed our existing warehouse and dashboards?

Yes. On-chain data is just another source: it lands in the same layered warehouse, under the same validation and modelling standards, and is read by Power BI, Looker Studio, Tableau or a Streamlit tool alongside everything else. Teams usually want it joined to off-chain records — customers, internal accounts, fiat movements — and that join is where most of the analytical value appears.

Your next step

Tell us what needs to work better.

Bring your goal, current tools and preferred working hours. We use the first 30-minute conversation to clarify fit and an initial scope, and you leave with a written next step.

Book a 30-minute call Email your brief contact@mvtech.solutions

Remote from Ahmedabad, India · Overlap with any timezone · No obligation