Careers

Join Yagoo. Put enterprise intelligence into real work.

We are looking for colleagues who want to turn data, knowledge, and processes into runnable capabilities.

Apply / inquire 199 0651 3493

Hangzhou · on-site collaboration

Open roles

Three teams, one mission

01

Algorithm Engineer

Responsibilities

  1. Turn existing raw/business data (documents, tables, system fields, logs, conversations, labeled data) into trainable, iterable domain datasets: clean → structure → label/weak-label → instructionize.
  2. Design training-data standards: instruction templates (single/multi-turn), tool-call formats, structured output schemas (JSON/table fields), refusal and boundary samples, hard cases and contrast sets.
  3. Build data-quality systems: dedup (semantic/fingerprint), noise filtering, sensitive-data handling, distribution stats, coverage analysis; produce measurable quality reports and improvements.
  4. Organize labeling and review (with business/ops/experts): guidelines, sampling rules, consistency checks; turn human experience into executable standards.
  5. Use models to speed work without trusting them blindly: rewriting, expansion, synthesis, first-pass labeling, hard-case mining, plus human review and alignment rules.
  6. Work with fine-tuning engineers to iterate data from training/eval feedback (fill gaps, add hard cases, fix templates, rebalance) so gains land on the data.
  7. Manage dataset assets with versioning, traceability, and rollback (data cards, change logs, samples, eval/control sets).

Requirements

  1. Bachelor’s or above; 2–5 years in data processing / NLP / ML engineering. End-to-end experience from raw data to trainable data is preferred.
  2. Data and engineering skills (~70%)
    • Strong Python (pandas/pyarrow/regex/json) and SQL; can write stable ETL, cleaning, and sampling scripts.
    • Text processing and QC: sentence split, token-length control, near-dup, anomaly detection, distribution stats.
    • Data versioning hygiene: naming, folders, metadata, change logs, reproducibility.
  3. LLM training-data literacy (~30%)
    • Understand instruction-tuning essentials: task definition, I/O boundaries, format constraints, multi-turn consistency, domain terms and factuality.
    • Can design eval/regression sets covering core tasks, boundaries, hard cases, and controls.
    • Know common alignment issues and data responses: hallucination control, refusal policy, style consistency, tool-call constraints.
  4. Communication and abstraction: turn experts’ spoken rules into trainable schemas, taxonomies, and sample rules; drive cross-team data delivery.

Nice to have

  1. Built labeling/QC systems or used mainstream annotation tools and workflows.
  2. Organized knowledge bases/graphs or enterprise field systems (business language → data language).
  3. Familiar with multimodal data organization for LLM training (if we expand later).
  4. Industry compliance/desensitization practice (audit trails, permission tiers, sensitive-field policy).

02

Operations Engineer

Responsibilities

  1. Operate, deploy, monitor, and stabilize the big-data platform, covering Hadoop / Spark / Flink / Kafka / Hive / HDFS / ClickHouse / Elasticsearch and related components for high availability, scale, and traceability.
  2. Build, upgrade, scale, migrate, and recover clusters, including Linux servers, storage, network, containers, and middleware.
  3. Build the ops system: alerting, logs, capacity planning, inspections, backup/restore, incident playbooks, permissions, configuration and change management.
  4. Tune performance and resources: compute, storage, scheduling, concurrency, IO, memory, and CPU.
  5. Work with development, algorithm, and data-engineering teams to keep ingest, cleaning, processing, analytics, and modeling running—deploy, monitor, repair, optimize.
  6. Strengthen security and compliance: accounts, access control, sensitive-data protection, log audit, vulnerability handling, and baseline hardening.
  7. Drive automation and DevOps: scripts and tools for deploy, inspection, alert response, log analysis, and recovery.

Requirements

  1. Bachelor’s or above, preferably in CS, software, information management, or networking; 3+ years of big-data platform ops, ideally on mid/large platforms or distributed clusters.
  2. Linux administration: servers, network, disks, processes, permissions, and system performance analysis.
  3. Hadoop ecosystem or mainstream big-data components: deploy, configure, operate, and troubleshoot HDFS, YARN, Hive, Spark, Flink, Kafka, ZooKeeper, and similar.
  4. Common middleware, databases, and observability: MySQL / PostgreSQL / Redis / Elasticsearch / Doris / Milvus / NebulaGraph / Prometheus / Grafana / ELK.
  5. Shell / Python scripting for automated inspection, batch ops, log analysis, and routine tooling.
  6. Understand distributed-system behavior; can handle node faults, service jitter, job failures, resource contention, network issues, disk alerts, and message backlog.
  7. Containers/virtualization preferred (Docker, Kubernetes); cloud or on-prem data-platform deployment is a plus.
  8. Strong incident analysis: find root cause quickly in complex environments and drive recovery.
  9. Collaborative and accountable; comfortable with fast-changing needs, long issue chains, and high stability bars.

Nice to have

  1. Large-cluster experience (100+ servers or PB-scale platforms).
  2. Kubernetes, DevOps, CI/CD, IaC (Ansible / Terraform).
  3. Enterprise data-governance practices: security, permissions, audit, desensitization, backup/restore.
  4. Ops experience on AI data platforms, knowledge bases, data middle platforms, or decision systems.
  5. Internal ops standards, platform standardization, or automation toolchain work.

03

Forward Deployed Engineer (FDE)

Responsibilities

  1. Scenario analysis and requirements: work on-site, understand industry processes and goals, identify valuable applications, turn objects/relations/events into implementable designs, and join validation and iteration.
  2. Data discovery and business understanding: analyze sources, structure, quality, and semantics. With AI-assisted analysis and FDE calibration, turn enterprise data into understandable, relatable business information for governance, ontology, and applications.
  3. Industry ontology and knowledge modeling: help turn industry knowledge, rules, and expert experience into machine-understandable models—entities, relations, events, rules, and knowledge links for agent reasoning.
  4. Agent and application development: design agent scenarios, compose Skills, orchestrate Workflows, configure tool calls, test and optimize so intelligence enters real processes.
  5. On-site delivery and continuous improvement: drive data access, validation, deployment, and rollout with customer teams; refine from usage feedback and turn ontology, Skills, Workflows, and agent templates into reusable assets.

Requirements

  1. Background in software, data engineering, or applied AI. Comfortable with Python, Java, and/or SQL; understand databases, APIs, data pipelines, and enterprise integration.
  2. Can read database structure, analyze relationships, spot quality issues, and turn technical data into business meaning.
  3. Familiar with LLMs, RAG, agents, workflows, knowledge graphs, and vector search; enterprise AI application experience preferred.
  4. Strong communication and problem analysis; work with customer business and tech teams to turn issues into technical plans.
  5. Fast learner who can enter an unfamiliar industry, complete understanding, analysis, design, and delivery.

Nice to have

  1. Experience in manufacturing digitalization, government IT, financial analysis, risk, data governance, or industry intelligence apps.
  2. Knowledge graphs, ontology modeling, data governance, LLM apps, agent apps, or RAG systems.
  3. Linux, Docker/Kubernetes, cloud deploy, distributed systems, data pipelines, or enterprise integration.
  4. Project delivery, pre-sales, solution design, or on-site customer collaboration.