Scope of work

Representation engineering

Inspect internal activations across layers, identify candidate directions associated with a target behavior, and test the association with control samples and interventions. A candidate direction is an observation about a particular model and dataset, not a mechanism assumed to transfer across models.

Refusal-direction analysis

In an authorized, controlled environment, analyze refusal patterns and how candidate directions, layer selection, and sampling choices affect them. Evaluation includes both benign requests refused in error and higher-risk requests that should still be refused; a single score is insufficient.

Abliteration and weight-level editing

Where justified, study directional ablation or other weight-level edits while preserving the original checkpoint, edit procedure, version history, and rollback route. Such edits can weaken safety behavior or degrade general capability. Production suitability depends on comparative evaluation and the customer's governance requirements.

Before-and-after benchmarking

Fix model versions, inference settings, and test data. Compare the target task, general capabilities, format adherence, false refusals, handling of higher-risk requests, stability, and performance. Report per-category results, failure cases, variation, and known limitations rather than using one demonstration prompt as proof.

Any experiment that changes safety behavior starts in an isolated environment with named access, permitted use, and release approval responsibilities. Public release or production deployment is outside the default experiment scope and requires a separate decision.

Start the conversation

Share the editable model version, the exact behavior to improve, redacted examples, and the capabilities and safety properties that must not regress. We will propose an experiment and comparison metrics before deciding whether to edit weights.

Related services: Model Post-Training · Private LLM Deployment