用 Jev 与 Laravel AI SDK 检测垃圾邮件和自动回复 ​

昨天,Taylor 宣布 Laravel AI SDK 的 1.x 分支已经支持 Jev。There There 当天便开始用它检测垃圾邮件。下面介绍 Jev 的工作方式及其用法。

Jev 与 LLM 有何不同 ​

Jev 由 TypeSafe 开发。LLM 会生成文本:向它提出问题,它会写出答案;若想获得结构化数据,还得在提示词中提出要求,并期待它按要求返回。Jev 则直接输出数值判断:只需向它提供状态信息和一个问题,它就会返回一个数字。

TypeSafe 将这类模型称为“系统一模型”(System One models)。它们像 LLM 一样读取自然语言,并从预先定义的答案中作出选择。其概率经过校准,也就是根据真实结果训练得出,因此在一批答案中,概率为 0.9 的判断应当约有九成正确。

实际使用时,既没有需要解析的自然语言回复,也无须用提示词要求模型返回有效 JSON。

Noul、Choice 与 Score ​

开发者可以使用三种问题类型定义可选答案。

Noul 是一道是非题。答案是一个数字,表示答案为“是”的概率。调用方式如下:

php
use Laravel\Ai\Classification;
use Laravel\Ai\Classification\Boolean;

$result = Classification::of('I have asked three times now. Can I please talk to a real person?')
    ->question('urgent', new Boolean('Does this request need an immediate response?'))
    ->classify();

$result['urgent']->probability;            // 0.94
$result['urgent']->isTrue(threshold: 0.8); // true

示例返回概率 0.94,随后可通过阈值判断得到布尔值 true。开发者可以自行决定阈值,因此每个问题都能设置不同的阈值。

Choice 会从开发者给定的一组选项中选择一个。示例如下:

php
use Laravel\Ai\Classification\Choice;

$result = Classification::of('My card was charged twice for order A-104. Please refund the duplicate.')
    ->question('department', new Choice('Which team should handle this request?', [
        'billing' => 'Payments, invoices, and refunds',
        'technical' => 'Bugs, outages, and integrations',
        'sales' => 'Pricing, plans, and upgrades',
    ]))
    ->classify();

$result['department']->choice;                   // 'billing'
$result['department']->probabilityOf('billing'); // 0.87
$result['department']->confidence;               // 0.82

除选中的选项外,返回结果还包含每个选项的概率,以及一个用于表示这些概率集中程度的置信度分数。

Score 会按照开发者自行描述的等级进行评分。例如,可以询问客户的不满程度,其中 0 表示平静,1 表示不满,2 表示非常愤怒。答案可以落在两个等级之间,因此 1.4 完全是一个有效结果。

状态无须是字符串。当一项决策取决于多个因素时,可以传入数组,为每个部分赋予名称:

php
Classification::of([
    'subject' => 'Duplicate charge',
    'message' => 'My card was charged twice for order A-104.',
    'order' => ['id' => 'A-104', 'charges' => [49, 49]],
    'refund_policy' => 'Duplicate charges are eligible for a refund.',
])->question('refund_due', new Boolean('The policy entitles this customer to a refund.'))
    ->classify();

虽然其中包含消息、订单和退款政策,但它依然是一个状态。

Jev 与其他提供商一样在 config/ai.php 中配置,并在环境变量文件中设置 TYPESAFE_API_KEY。

要解决的问题 ​

同样在昨天,团队发布了 There There,这是一款新的帮助台产品。帮助台收到的许多邮件并非来自客户,例如外出自动回复、退信、订阅确认和 DMARC 报告。这些邮件不应进入收件箱,也不值得花费 LLM 调用成本为每一封生成标题和摘要。

有些邮件会在邮件头中表明自身类型,例如 Auto-Submitted、空返回路径,或名为 mailer-daemon 的发件人。检查这些信息没有成本,因此系统会优先检查邮件头。

许多邮件服务器并不会设置这些邮件头。为此,系统曾维护一份邮件主题前缀列表,收录了十五种语言中的 Automatische Antwort、Réponse automatique、Out of office 等表达;退信另有九个前缀。此外,系统还设置了规则,避免把询问自动回复相关问题的客户邮件误判为自动回复。

每当发现列表遗漏的邮件,就需要再向其中添加一个字符串。

用 Jev 替换关键词列表 ​

邮件头检查依然最先执行,所有无法据此判断的邮件都会交给 Jev。

每个问题只描述一次:将它定义为枚举中的一个 case,并让它携带自身的阈值。这样,添加第四个问题时,原本涉及四个文件的修改便可收敛为新增一个 case。

php
enum InboundJudgement: string
{
    case IsAutoResponse = 'is_auto_response';
    case IsBounce = 'is_bounce';
    case IsSpam = 'is_spam';

    public function question(): Boolean
    {
        return match ($this) {
            self::IsAutoResponse => new Boolean(
                'A system sent this mail on its own, rather than a person choosing to write to us.',
                [
                    'true' => 'Sent on a trigger with no human involved at send time: out-of-office
                        notices, delivery reports, ticket acknowledgements, subscription
                        confirmations, digests and alerts. Wording composed in advance still counts',
                    'false' => 'A person sat down and sent this. Still false when a contact form or
                        chat widget wrapped their words in a template and added lines such as Name,
                        E-mail or Subject',
                ],
            ),
            self::IsSpam => new Boolean(
                'This mail is unsolicited bulk mail, a scam, or phishing rather than a genuine
                    message from a customer.',
                [
                    'true' => 'Cold sales outreach, marketing blasts, scams, phishing, or anything
                        the recipient never asked for',
                    'false' => 'A real person writing about the product, their account, or their own
                        support request, however brief or badly written',
                ],
            ),
            // ...
        };
    }

    public function threshold(): float
    {
        return match ($this) {
            self::IsAutoResponse => 0.75,
            self::IsBounce, self::IsSpam => 0.9,
        };
    }
}

其中 true 和 false 的描述属于可选项,但建议补充。它们对判断的指导作用比上方的问题文本本身更大。

所有无法通过邮件头判断的问题都会在一次请求中发送。Jev 只读取一次状态,再并行回答这些问题;由于只对输入 token 收费,同时询问三个问题与询问一个问题的成本相同。执行这项工作的 Action 如下:

php
public function execute(Message $message, Ticket $ticket, Workspace $workspace): void
{
    $judgements = array_filter(
        InboundJudgement::cases(),
        fn (InboundJudgement $judgement) => ! $judgement->settledByHeaders($message),
    );

    if ($judgements === []) {
        return;
    }

    try {
        $response = Classification::of([
            'subject' => $ticket->subject,
            'from_name' => $message->author_name,
            'from_email' => $message->author_email ?? $ticket->contact?->email,
            'message' => Str::limit($message->body_text, 10_000),
        ])
            ->questions($this->questionsFor($judgements))
            ->timeout(10)
            ->classify();

        $verdicts = $this->verdicts($judgements, $response);
    } catch (Throwable $exception) {
        Log::warning('Could not classify an inbound message.', [
            'message_id' => $message->id,
            'error' => $exception->getMessage(),
        ]);

        return;
    }

    $message->updateQuietly([...$verdicts, 'classification' => $response->answers]);
}

其中有两处做法值得借鉴。整个调用都位于 try 块内;发生问题时记录日志并结束本次分类,让邮件处理流程继续运行,因为分类只是增强功能,不应中断它所辅助的邮件处理流水线。系统还会截断邮件,因为入站邮件可能长达数 MB,而超出开头部分的内容不会改变邮件所属的类别。

把答案转换为布尔值时,会应用各个问题自己的阈值。系统会同时保存原始概率,以便日后调整阈值,并查看新阈值原本会产生什么结果:

php
private function verdicts(array $judgements, ClassificationResponse $response): array
{
    $verdicts = [];

    foreach ($judgements as $judgement) {
        $answer = $response->answer($judgement->value);

        $verdicts[$judgement->value] = $answer->isTrue($judgement->threshold());
    }

    return $verdicts;
}

这些判断结果会存储在消息上。当客户在 There There 中创建包含“Is spam”条件的工作流时,检查该条件只需读取一列数据,完全不会调用 Jev。

使用体验 ​

Jev 专注于一件小事:给出一个数字,其余决策则留在开发者自己的代码中,便于阅读和测试。

它的响应足够快,成本也足够低,使用时几乎无须额外权衡。一次同时回答三个问题的邮件分类耗时 639 毫秒,并行运行时每秒可处理约 48 封邮件。团队没有花时间专门调优,因此仍有进一步提升空间;对于当前用途,这样的速度已经足够。Jev 每百万输入 token 的费用为 0.042 美元,输出 token 免费。按 There There 的使用情况计算,每封邮件的成本约为 0.04 美分,每月约为 0.36 美元。

There There 以及团队的其他产品中,还有许多计划使用 Jev 的场景。更多由 Jev 驱动的功能很快就会推出。

如需进一步了解,可以查阅 TypeSafe 文档和 Laravel AI SDK。如需实际体验垃圾邮件检测功能,可以试用 There There。

本作品采用《CC 协议》,转载必须注明作者和本文链接