Security and fraud

The call that should never have gone through – second by second

Last week's attack on Wall Street's largest hedge funds did not come through the payment flow. It came through the service desk, in a voice the employees recognised.

IngridIngrid,
Close-up of a hand hesitating over the handset of a desk phone while the display shows an incoming number, in a blurred office

On Wednesday 5 August, Bloomberg reported that several of the world's largest hedge funds – among them Citadel, Millennium Management, Two Sigma and Point72 – were targeted in a coordinated voice phishing campaign. The attackers used AI-generated voices in phone calls and voice messages to get employees to hand over login credentials or approve access to internal systems. Two Sigma says the attempt was blocked with no impact to its data or systems. Point72 notified its investors and found no indication that client data had been taken. Citadel and Millennium have not commented publicly.

The interesting part is not who was called. It is where the attack landed: not in the payment flow, but in the service desk. The inbound phone.

Below is that call broken into its parts. Not because your company is a hedge fund, but because every one of those parts exists in an ordinary service desk – and because several of them are easier to close than people think.

Second 0: the number looks right

The call arrives from a number that looks internal, or that is already in the contact list. Making that happen costs very little.

Caller ID is not a check. It is a piece of information that travels with the call, and in practice the caller chooses it. It is still the first signal the person answering responds to, and it colours everything that follows.

What has to be true at this point: nobody in the organisation believes a number proves anything.

Second 6: the voice is familiar

Then the person speaks. It is the operations manager. Or IT support. Or a colleague from another office you have heard on a video call a handful of times.

A few seconds of public audio is enough to make a usable copy – a webinar, a podcast, a recorded voicemail greeting. Recognition is not a check. It is a feeling, and it is now trivial to produce. That the voice is a biometric characteristic has consequences well beyond fraud, but the practical consequence is this: the voice can no longer be what decides who you are talking to.

Second 25: the request is boring, and therefore believable

The attacker does not ask for a wire transfer. He says he has switched phones and cannot receive the one-time code. Or that he is at the airport and locked out of the VPN.

This is the part people underestimate. A demand for a large sum triggers suspicion. A password reset does not, because it is the single most common request a service desk handles. The attack hides inside the ordinary.

Second 40: the pressure arrives, not the question

There is always a reason it is urgent, and the reason is always reasonable: a board meeting starts in ten minutes, a client is waiting, a deadline falls today.

Time pressure is not in itself a sign of fraud. But it is the most consistent common denominator in these calls, because it is the time pressure that makes the person answering skip the step she would otherwise have taken.

Second 55: the request is small enough to say yes to

Read out the code. Approve the prompt on your phone. Reset the password. Add this number as an authentication device.

None of it looks like giving away access. Every one of them is.

Where the call can actually be stopped

Everything above is description. This is the point: there is only one place where this call can be stopped reliably, and it is where the check cannot be supplied by the caller.

A one-time code read out loud is not a check – the caller supplies it himself. A question about a date of birth or an employee number is not a check – the answer sits in data that has already leaked in several places. The voice sounding right is not a check.

What is a check: hanging up and calling back on the number registered in the HR system. Or asking a manager for approval through a channel other than the one the call came in on. It is not elegant, and it takes two minutes.

The hard part is organisational, not technical: the person answering has to be allowed to spend those two minutes on someone who sounds like the boss, and has to know that nobody will criticise her for it afterwards. It has to be written down, and it has to be said out loud. Otherwise ordinary politeness beats the procedure every time.

Afterwards: what the log has to be able to show

Two Sigma was able to say the attempt was blocked. To say something like that, you have to know what actually happened in the calls.

Three questions are worth being able to answer tomorrow: which calls asked for a change in access, what was decided, and who decided it. If the answer only exists in the memory of the person who was on the phone, it does not exist.

That applies just as much when a voice agent takes the call. An agent can do this better than a person – it does not get flustered, it does not skip steps to be pleasant, and it can log every stage. It can also do it worse, if it has been set up to be helpful without limits. What the call logs look like, and what the agent is allowed to do without confirmation from another system, belongs in the assessment of the vendor – not in the first operations review.

The voice is a channel, not an identity document

It is worth keeping the two apart. A compliant voice agent says in its opening line that it is AI. An attacker using a synthetic voice never does. He uses it to sound like a human being you know.

The difference is not in the technology. It is in whether the call can withstand being accounted for afterwards.

Threll.ai builds voice agents in Norwegian, Swedish and Danish, on Nordic telecoms infrastructure. What makes a call safe is not that it sounds right. It is that you can show what was said, what was done, and why.

Related Articles

Continue reading more articles