Discover your dream Career
For Recruiters

How hedge fund Balyasny's AI team is performing better than OpenAI

One of the most important tools for dealing with large lancuage models (LLMs) is retrieval augmented generation (RAG), the process of searching for data outside a model's training data to better answer a question. Using this technique in finance can be problematic... unless you work for hedge fund Balyasny.

Click here to follow our new WhatsApp channel, and get instant news updates straight to your phone 📱

Balyasny has been building out its applied AI research team with Google and DeepMind alums to create its own AI tools, like AI chatbot BAMChatGPT. The team is led by Charlie Flanagan, an ex-Google data scientist, and its tools are used by 80% of the fund's employees according to lead engineer Michal Mucha. Business Insider says the general purpose of Balyasny's various tools is to answer complex questions for traders about certain stocks and gauge the impact of global events on their portfolios.

In an academic paper earlier this month, the hedge fund announced a new creation; BAM Embeddings. Embeddings are a key part of the RAG process, helping its LLM contextualize outside data, and Balyasny's creation has been made specifically for the "arcane jargon" of financial services. BAMChatGPT is already more specialised for finance than general purpose chatbots, but RAG will be a necessity for certain queries, like asking the bot about recent news stories or finanical results outisde its training data; general-purpose embeddings may limit the LLM's accuracy when accessing external sources as those embeddings may not properly understand financial jargon and not pick the relevant data to analyze.

Early indications are that the embeddings are working. When testing how often the LLM (in this case, the Mistral 7B Instruct model) would return the most relevant passage in a dataset of financial documents, BAM Embeddings did so over 60% of the time, while OpenAI's own general purpose embeddings did so less than 40%.

The paper also tested on publicly available LLM performance benchmarking system FinanceBench. There, BAM Embeddings answered queries with 55% accuracy, compared to the 47% of OpenAI's ada-002 embeddings.

It's not perfect yet, however. The FinanceBench results showed queries using BAM Embeddings were incorrect 30% of the time. LLMs are prone to hallucinating, even with RAG; A study in July revealed that OpenAI's GPT-3.5, thought to be the most popular LLM in an algorithmic trading context, hallucinated in 27.8% of responses in which RAG was used. On top of that, BAM Embeddings was "trained on a dataset of 14.3M synthetic queries", which some believe increases the risk of hallucination.

Nonetheless, it seemingly gives Balyasny a competitive advantage in the increasingly crowded field of people using RAG. A report from VC firm Menlo Ventures says that RAG is the primary architectural approach in AI for 51% of enterprises in 2024, up from 31% the year before.

Have a confidential story, tip, or comment you’d like to share? Contact: Telegram: @AlexMcMurray, WhatsApp: (+1 269 237 3950)Click here to fill in our anonymous form, or email editortips@efinancialcareers.com.

Bear with us if you leave a comment at the bottom of this article: all our comments are moderated by human beings. Sometimes these humans might be asleep, or away from their desks, so it may take a while for your comment to appear. Eventually it will – unless it’s offensive or libelous (in which case it won’t.)

author-card-avatar
AUTHORAlex McMurray Reporter

Sign up to Morning Coffee!

Coffee mug

The essential daily roundup of news and analysis read by everyone from senior bankers and traders to new recruits.

Sign up to Morning Coffee!

Coffee mug

The essential daily roundup of news and analysis read by everyone from senior bankers and traders to new recruits.