Introducing the Open FinLLM Leaderboard

The growing complexity of financial language models (LLMs) necessitates evaluations that go beyond general NLP benchmarks. While traditional leaderboards focus on broader NLP tasks like translation or summarization, they often fall short in addressing the specific needs of the finance industry. Financial tasks, such as predicting stock movements, assessing credit risks, and extracting information from financial reports, present unique challenges that require models with specialized skills. This is why we decided to create the Open FinLLM Leaderboard.

Content Credentials

Critical information about the content you see online is often inaccessible or inaccurate. Content Credentials are a new open technology for revealing answers to your questions about content with a simple click: How was it made? Is it AI-generated? When was it created or edited?

Lawsuit Argues Warrantless Use of Flock Surveillance Cameras Is Unconstitutional

“The City of Norfolk, Virginia, has installed a network of cameras that make it functionally impossible for people to drive anywhere without having their movements tracked, photographed, and stored in an AI-assisted database that enables the warrantless surveillance of their every move. This civil rights lawsuit seeks to end this dragnet surveillance program,” the lawsuit notes. “In Norfolk, no one can escape the government’s 172 unblinking eyes,” it continues, referring to the 172 Flock cameras currently operational in Norfolk. The Fourth Amendment protects against unreasonable searches and seizures and has been ruled in many cases to protect against warrantless government surveillance, and the lawsuit specifically says Norfolk’s installation violates that. 

NVIDIA Unveils “Industry Leading” Open-Source Llama-3.1-Nemotron-70B-Instruct LLM , Surpassing OpenAI’s GPT-4o In AI-Focused Benchmarks

Interestingly, based on the Llama-3.1-Nemotron-70B-Instruct LLM model card present at HuggingFace, this particular model manages to solve the “strawberry” problem, which traditional AI models were unable to solve, where it involved counting the R’s in the word. This isn’t just the only achievement, as the upcoming details might surprise readers more. NVIDIA’s Llama-3.1-Nemotron-70B-Instruct LLM has achieved leading ranking at numerous benchmarks, notably Arena Hard, an automatic evaluation tool for instruction-tuned LLMs, and here’s how the overall scores stack up.

Everything I built with Claude Artifacts this week

I’m a huge fan of Claude’s Artifacts feature, which lets you prompt Claude to create an interactive Single Page App (using HTML, CSS and JavaScript) and then view the result directly in the Claude interface, iterating on it further with the bot and then, if you like, copying out the resulting code.

Dimensions of AI Literacies – Opened Culture

Another set of literacies. I like the structure but wonder how many different sets of literacies we have at this point- multimodal, computational thinking, digital, etc. etc. There is likely a point where you have to have things that apply broadly enough that you don’t have to rethink it with each new evolution of technology.

h/t Sarah LW