Mobile application behavior datasets
Source-level app behaviour, permissions, endpoints, consent timing, and observed data flows.
Data Solutions
The same discipline behind our systems and public evidence, delivered as data and infrastructure: a data warehouse, reproducible environments, expert trajectories, and evaluations for difficult, long-horizon work. Our Rust and Axum services connect through a SQL bridge to ternary-logic pipelines. Your corpus stays private. Every delivered artefact stays traceable.
Source-level app behaviour, permissions, endpoints, consent timing, and observed data flows.
Reusable SDKs, tracker relationships, infrastructure, and cross-application recurrence mapped as connected evidence.
Tool-use traces under injection, conflicting instructions, boundary pressure, recovery, and stop conditions.
Multi-domain observations, entity relationships, change histories, and evidence-linked causal chains.
Controlled, reproducible perturbations of open benchmarks for robustness testing beyond clean inputs.
What we deliver
We define success criteria with your team, then build the data operation around the way your models are actually trained and evaluated. No anonymous task stream and no benchmark theatre: the environment, trace, label, and decision remain connected.
A governed foundation for operational, training, and evaluation data, connected to our Rust/Axum services through a production SQL bridge.
Human-simulated companies, computer-use and MCU mockups, deterministic resets, and controlled credentials for repeatable agent work.
MCP-bench and TAU-bench extensions, TinyTAU for on-device agents, and task suites built around your real tools and constraints.
Expert demonstrations, step-level annotations, preference labels, failure taxonomies, and calibrated evaluations for training and reward shaping.
Coding-agent safety, MCP injection assessment, computer-use injection red teaming, and adversarial trajectories with reproducible evidence.
Repository-scale generation, issue resolution, code review, testing, debugging, tool-use, data analysis, and visual frontend tasks.
Built for
Task environments and data shaped around how each agent observes, reasons, uses tools, and recovers when the work stops going to plan.
Natural-language dialogue grounded in evidence and policy.
Internal tools, workflows, knowledge bases, and governed automation.
Multi-source investigation, synthesis, and defensible conclusions.
Browsers, applications, filesystems, and realistic interfaces.
Code writing, debugging, repository work, testing, and review.
Desktop, mobile, embedded, wearable, and on-device environments.
How it works
Objectives, constraints, rubrics, formats, and acceptance criteria become a testable specification.
We create the environment and tasks. Experts perform the work while the complete raw trajectory is recorded.
Automated checks enforce schemas, invariants, logical consistency, rubric adherence, and task completion.
Senior reviewers audit difficult and flagged traces plus a statistically meaningful sample of the remainder.
Versioned datasets, evaluation reports, deterministic environments, and audit logs arrive ready for your workflow.
Three ways to start
Use an already produced and validated dataset. We align the format and delivery boundary with your stack.
Start a prepared collection or evaluation pipeline, adjusted to your volume, domain, and acceptance criteria.
Extend an existing corpus, adapt it to your domain, or commission a new environment and dataset from first principles.
Expert depth
The work is produced and reviewed at the level where our own systems are built: Python, Rust, C/C++, JavaScript and TypeScript, Go, Java, Kotlin, SQL, Bash, mobile and ML stacks — across backend, frontend, systems, security, DevOps, data science, and model engineering.
this is a useless cookie banner. it's just here to look like one * we don't use cookies, so there's nothing to consent to. don't let anyone tell you otherwise.two buttons, one closes this and throws some confetti. the other literally does nothing.