Weekly Web Harvest for 2024-10-20

  • Beyond Tools: LLMs and the Emergence of Extended Cognition | Psychology Today
    This mirroring effect is transformative in ways we’re only beginning to understand. When we engage with an LLM, we’re compelled to externalize our internal thought processes, making them more visible and, therefore, more amenable to refinement. Like a skillful conversation partner, the system prompts us to clarify our assumptions and elaborate on our logic, creating a feedback loop that leads to deeper understanding.

  • The case for human–AI interaction as system 0 thinking | Nature Human Behaviour
    “The term system 0 is chosen deliberately to emphasize its foundational and pervasive role in modern cognition. Unlike the system 1 and system 2 (which operate within the individual mind), system 0 forms an artificial, non-biological underlying layer of distributed intelligence that interacts with and augments both intuitive and analytical thinking processes. This designation underscores its function as a preprocessor and enhancer of information, which actively shapes the inputs to traditional cognitive systems rather than simply extending them.”

  • diegogalpy/sChemNET
    sChemNET: A deep learning framework for predicting small molecules targeting microRNA function.
  • microsoft/RD-Agent: Research and development (R&D) is crucial for the enhancement of industrial productivity, especially in the AI era, where the core aspects of R&D are mainly focused on data and models. We are committed to automating these high-value ge
    RDAgent aims to automate the most critical and valuable aspects of the industrial R&D process, and we begin with focusing on the data-driven scenarios to streamline the development of models and data. Methodologically, we have identified a framework with two key components: ‘R’ for proposing new ideas and ‘D’ for implementing them. We believe that the automatic evolution of R&D will lead to solutions of significant industrial value.

  • [2409.16191] HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
    Besides, we propose Hierarchical Long Text Evaluation (HelloEval), a human-aligned evaluation method that significantly reduces the time and effort required for human evaluation while maintaining a high correlation with human evaluation. We have conducted extensive experiments across around 30 mainstream LLMs and observed that the current LLMs lack long text generation capabilities. Specifically, first, regardless of whether the instructions include explicit or implicit length constraints, we observe that most LLMs cannot generate text that is longer than 4000 words. Second, we observe that while some LLMs can generate longer text, many issues exist (e.g., severe repetition and quality degradation). Third, to demonstrate the effectiveness of HelloEval, we compare HelloEval with traditional metrics (e.g., ROUGE, BLEU, etc.) and LLM-as-a-Judge methods, which show that HelloEval has the highest correlation with human evaluation.
  • Project CETI •– Home
    CETI is a nonprofit organization applying advanced machine learning and state-of-the-art robotics to listen to and translate the communication of sperm whales. Our research focus is in Dominica in the Eastern Caribbean.

  • AI Hackathon | TEDAI San Francisco
  • Asana AI for Work & Project Management • Asana
  • Runway Research | Introducing Act-One
  • Roblox: Inflated Key Metrics For Wall Street And A Pedophile Hellscape For Kids – Hindenburg Research
    The title sums it up pretty well.
  • GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models – Apple Machine Learning Research
    Furthermore, we investigate the fragility of mathematical reasoning in these models and show that their performance significantly deteriorates as the number of clauses in a question increases. We hypothesize that this decline is because current LLMs cannot perform genuine logical reasoning; they replicate reasoning steps from their training data. Adding a single clause that seems relevant to the question causes significant performance drops (up to 65%) across all state-of-the-art models, even though the clause doesn’t contribute to the reasoning chain needed for the final answer.
  • Liquid Neural Networks | Ramin Hasani | TEDxMIT – YouTube
    A good explanation of LLM limitations in addition to a possible future path with liquid neural networks
  • Luddites Win
    It’s all supposed to be some sort of “life hack” – except nothing is getting easier or better; we’re not happier. Everything’s just getting more hackneyed, more hurried, more chopped up and garbled and shredded. We can feel it. Everyday everything is more and more fragile, more and more precarious.
  • Interview series on risks from AI – LessWrong
    In 2011, Alexander Kruel (XiXiDu) started a Q&A style interview series asking various people about their perception of artificial intelligence and possible risks associated with it.

  • HarmonyCloak
    By embedding imperceptible, error-minimizing noise into the music, HarmonyCloak effectively prevents AI systems from extracting meaningful patterns, all while preserving the perceptual quality of the music for human listeners.
  • AI Snake Oil: What Artificial Intelligence Can Do, What It Can’t, and How to Tell the Difference: Narayanan, Arvind, Kapoor, Sayash: 9780691249131: Amazon.com: Books
  • Claude | Computer use for coding – YouTube
    IT’s weird to think about the efficiency of the process with stuff like this. We build the GUI, mouse, etc. to let humans avoid the code and interact with computers. Then, with this, we have computers translate those same things . . . to interact with computers. It seems interesting but misguided.
  • Introducing the Open FinLLM Leaderboard
    The growing complexity of financial language models (LLMs) necessitates evaluations that go beyond general NLP benchmarks. While traditional leaderboards focus on broader NLP tasks like translation or summarization, they often fall short in addressing the specific needs of the finance industry. Financial tasks, such as predicting stock movements, assessing credit risks, and extracting information from financial reports, present unique challenges that require models with specialized skills. This is why we decided to create the Open FinLLM Leaderboard.
  • Content Credentials
    Critical information about the content you see online is often inaccessible or inaccurate. Content Credentials are a new open technology for revealing answers to your questions about content with a simple click: How was it made? Is it AI-generated? When was it created or edited?

  • Lawsuit Argues Warrantless Use of Flock Surveillance Cameras Is Unconstitutional
    “The City of Norfolk, Virginia, has installed a network of cameras that make it functionally impossible for people to drive anywhere without having their movements tracked, photographed, and stored in an AI-assisted database that enables the warrantless surveillance of their every move. This civil rights lawsuit seeks to end this dragnet surveillance program,” the lawsuit notes. “In Norfolk, no one can escape the government’s 172 unblinking eyes,” it continues, referring to the 172 Flock cameras currently operational in Norfolk. The Fourth Amendment protects against unreasonable searches and seizures and has been ruled in many cases to protect against warrantless government surveillance, and the lawsuit specifically says Norfolk’s installation violates that. 

  • NVIDIA Unveils “Industry Leading” Open-Source Llama-3.1-Nemotron-70B-Instruct LLM , Surpassing OpenAI’s GPT-4o In AI-Focused Benchmarks
    Interestingly, based on the Llama-3.1-Nemotron-70B-Instruct LLM model card present at HuggingFace, this particular model manages to solve the “strawberry” problem, which traditional AI models were unable to solve, where it involved counting the R’s in the word. This isn’t just the only achievement, as the upcoming details might surprise readers more. NVIDIA’s Llama-3.1-Nemotron-70B-Instruct LLM has achieved leading ranking at numerous benchmarks, notably Arena Hard, an automatic evaluation tool for instruction-tuned LLMs, and here’s how the overall scores stack up.

  • Spirit LM – Interleaved Spoken and Written Language Model
    We introduce Spirit LM, a foundation multimodal language model that freely mixes text and speech. Our model is based on a 7B pretrained text language model that we extend to the speech modality by continuously training it on text and speech units.
  • Everything I built with Claude Artifacts this week
    I’m a huge fan of Claude’s Artifacts feature, which lets you prompt Claude to create an interactive Single Page App (using HTML, CSS and JavaScript) and then view the result directly in the Claude interface, iterating on it further with the bot and then, if you like, copying out the resulting code.

  • X changed its terms of service to let its AI train on everyone’s posts. Now users are up in arms | CNN Business
  • Dimensions of AI Literacies – Opened Culture
    Another set of literacies. I like the structure but wonder how many different sets of literacies we have at this point- multimodal, computational thinking, digital, etc. etc. There is likely a point where you have to have things that apply broadly enough that you don’t have to rethink it with each new evolution of technology.

    h/t Sarah LW

  • Do AI Detectors Work? Students Face False Cheating Accusations – Bloomberg
    The students most susceptible to inaccurate accusations are likely those who write in a more generic manner, either because they’re neurodivergent like Olmsted, speak English as a second language (ESL) or simply learned to use more straightforward vocabulary and a mechanical style, according to students, academics and AI developers. A 2023 study by Stanford University researchers found that AI detectors were “near-perfect” when checking essays written by US-born eighth grade students, yet they flagged more than half of the essays written by nonnative English students as AI-generated. OpenAI recently said it has refrained from releasing an AI writing detection tool in part over concerns it could negatively affect certain groups, including ESL students.

  • Install – Display Posts
    I am no longer comfortable hosting my code on WordPress.org given recent actions by Matt Mullenweg:

  • Philly artists with disabilities shown at the Painted Bride – WHYY
    Created by a collective of artists and activists, the Undue Burden archive catalogs the creative lives of disabled, neurodivergent and chronically ill Philadelphians.