World ModelData SolutionsEvidenceAccess

Data Solutions

production, data and agents that have to work

The same discipline behind our systems and public evidence, delivered as data and infrastructure: a data warehouse, reproducible environments, expert trajectories, and evaluations for difficult, long-horizon work. Our Rust and Axum services connect through a SQL bridge to ternary-logic pipelines. Your corpus stays private. Every delivered artefact stays traceable.

Mobile application behavior datasets

Source-level app behaviour, permissions, endpoints, consent timing, and observed data flows.

SDK & tracker knowledge graphs

Reusable SDKs, tracker relationships, infrastructure, and cross-application recurrence mapped as connected evidence.

Agent safety trajectories

Tool-use traces under injection, conflicting instructions, boundary pressure, recovery, and stop conditions.

World-state & causality mapping

Multi-domain observations, entity relationships, change histories, and evidence-linked causal chains.

Stressed benchmark variants

Controlled, reproducible perturbations of open benchmarks for robustness testing beyond clean inputs.

What we deliver

From environment to evaluation

We define success criteria with your team, then build the data operation around the way your models are actually trained and evaluated. No anonymous task stream and no benchmark theatre: the environment, trace, label, and decision remain connected.

01

Data warehouse

A governed foundation for operational, training, and evaluation data, connected to our Rust/Axum services through a production SQL bridge.

02

Virtual environments

Human-simulated companies, computer-use and MCU mockups, deterministic resets, and controlled credentials for repeatable agent work.

03

Capability evaluations

MCP-bench and TAU-bench extensions, TinyTAU for on-device agents, and task suites built around your real tools and constraints.

04

Trajectory data

Expert demonstrations, step-level annotations, preference labels, failure taxonomies, and calibrated evaluations for training and reward shaping.

05

Safety evaluation

Coding-agent safety, MCP injection assessment, computer-use injection red teaming, and adversarial trajectories with reproducible evidence.

06

Coding data

Repository-scale generation, issue resolution, code review, testing, debugging, tool-use, data analysis, and visual frontend tasks.

Built for

Agents across the whole tool chain

Task environments and data shaped around how each agent observes, reasons, uses tools, and recovers when the work stops going to plan.

Conversational agents

Natural-language dialogue grounded in evidence and policy.

Corporate assistants

Internal tools, workflows, knowledge bases, and governed automation.

Deep-research agents

Multi-source investigation, synthesis, and defensible conclusions.

Computer-use agents

Browsers, applications, filesystems, and realistic interfaces.

Coding copilots

Code writing, debugging, repository work, testing, and review.

OS & device agents

Desktop, mobile, embedded, wearable, and on-device environments.

How it works

A managed pipeline, built by engineers for engineers

01

Define

Objectives, constraints, rubrics, formats, and acceptance criteria become a testable specification.

02

Build & execute

We create the environment and tasks. Experts perform the work while the complete raw trajectory is recorded.

03

Validate

Automated checks enforce schemas, invariants, logical consistency, rubric adherence, and task completion.

04

Review

Senior reviewers audit difficult and flagged traces plus a statistically meaningful sample of the remainder.

05

Deliver

Versioned datasets, evaluation reports, deterministic environments, and audit logs arrive ready for your workflow.

Three ways to start

Data that meets you where the work is

01

Ready to deliver

Use an already produced and validated dataset. We align the format and delivery boundary with your stack.

02

Pipeline-ready

Start a prepared collection or evaluation pipeline, adjusted to your volume, domain, and acceptance criteria.

03

Expand or customize

Extend an existing corpus, adapt it to your domain, or commission a new environment and dataset from first principles.

Expert depth

Not crowd work. Technical work.

The work is produced and reviewed at the level where our own systems are built: Python, Rust, C/C++, JavaScript and TypeScript, Go, Java, Kotlin, SQL, Bash, mobile and ML stacks — across backend, frontend, systems, security, DevOps, data science, and model engineering.

© 2026 RFI-IRFOS  ·  Graz, Austria  ·  GISA 39261441  ·  UID ATU83405245

Human rights are not subject to negotiation.

Foundation

Intelligence

World Model

Applications

Repositories

Crates (crates.io)

all packages →

this is a useless cookie banner. it's just here to look like one * we don't use cookies, so there's nothing to consent to. don't let anyone tell you otherwise.two buttons, one closes this and throws some confetti. the other literally does nothing.