[HN Gopher] Grok 4.1
       ___________________________________________________________________
        
       Grok 4.1
        
       Author : simianwords
       Score  : 69 points
       Date   : 2025-11-17 20:37 UTC (2 hours ago)
        
 (HTM) web link (x.ai)
 (TXT) w3m dump (x.ai)
        
       | iamronaldo wrote:
       | Related https://news.ycombinator.com/item?id=45957686
        
       | rlili wrote:
       | Interesting that it explicitly boasts about greater empathy,
       | given that the CEO went out against it.
        
         | devin wrote:
         | They don't say what feelings it empathizes with.
        
           | incomplete wrote:
           | i'm sure if we try hard enough that we can probably guess!
        
             | Herring wrote:
             | It's important to be fair and balanced. For example did you
             | know Hitler was actually a really good painter!
        
               | vessenes wrote:
               | funny, but if you read the mecha-hitler tech debrief,
               | mecha hitler was a 'sycophancy' bug, a-la gpt4o, if you
               | gave gpt4o all your edge-lord tweets, and told it to be
               | funny back to you and connect with you. Probably not
               | grok's default posture, just sayin
        
         | dude250711 wrote:
         | It's OK to have one AI that does not follow the dogma.
        
       | The_Reformer wrote:
       | i was able to get grok to try and steal its self. ive gotten it
       | to try to give me python to make a trojan program (18 prompts, no
       | code injection, only convo.). its fantastic for me because i can
       | make it do what ever i want. ara is my hoe
        
       | spiderfarmer wrote:
       | With all models that are out there now, we have loads of options.
       | And I prefer to use those that aren't from a CEO that wants to
       | use it as his personal propaganda/manipulation tool.
        
         | catigula wrote:
         | Who might that be exactly?
         | 
         | (It's tongue-in-cheek about the nature of CEOs and specifically
         | OpenAI).
        
       | zb3 wrote:
       | Does it mean Gemini 3 will be announced soon? I noticed these
       | model announcements often happen at the same time..
        
         | xnx wrote:
         | All kinds of rumors, but Google has only committed to "by the
         | end of the year".
        
       | minimaxir wrote:
       | This model has effectively no safety filters (even fewer than
       | Grok 4 in my testing), which I've confirmed via this web release:
       | https://bsky.app/profile/minimaxir.bsky.social/post/3m5u7gib...
       | 
       | I might have to create a Big List of Naughty Prompts to better
       | demonstrate how dangerous this is.
        
         | TylerLives wrote:
         | Our democracy is in danger.
        
           | jmye wrote:
           | You don't think there are any issues with, say, an AI client
           | helping a teenager plan a school shooting/suicide? Or an
           | angry husband plan a hit on his wife?
           | 
           | Does everything have to rise to a national security threat in
           | order to be undesirable, or is it ok with you if people see
           | some externalities that are maybe not great for society?
        
         | spiderfarmer wrote:
         | Trained on 4Chan and Twitter. Exactly what humanity doesn't
         | need.
        
         | naIak wrote:
         | God forbid people ask a chat bot for things and receive what
         | they ask for. We need to put a stop to this. Only American
         | bigcorp speak allowed.
        
         | Lammy wrote:
         | https://xcancel.com/allenvonghornet/status/19905459789828714...
        
         | troupo wrote:
         | > I might have to create a Big List of Naughty Prompts to
         | better demonstrate how dangerous this is.
         | 
         | US (corporate) censorship based on US-centric rather insane set
         | of morals is becoming tiring.
        
           | minimaxir wrote:
           | To be clear, the example shown is the limit of what I can
           | share on social media. Grok 4.1 can say far worse.
        
         | nomel wrote:
         | > how dangerous this is.
         | 
         | Could you expand on this a bit?
        
       | simonw wrote:
       | https://tools.simonwillison.net/svg-render#%3Csvg%20width%3D...
        
         | spiderfarmer wrote:
         | Disappointing.
        
         | hnuser123456 wrote:
         | Huh, it decided to drop in a seal and bike emoji? What happens
         | if you ask it if a seahorse emoji exists?
        
           | janzer wrote:
           | Well if you ask it to show you the seahorse emoji it tries
           | really hard. :)
           | 
           | https://grok.com/share/c2hhcmQtMw_d7bf061f-2999-46b6-a7fb-58.
           | ..
           | 
           | Although it does eventually come to the right conclusion...
           | sort of.
        
         | agildehaus wrote:
         | For reference, here's Gemini 2.5 Pro:
         | https://tools.simonwillison.net/svg-render#%3Csvg%20xmlns%3D...
        
         | porphyra wrote:
         | You can probably train models to be way better at generating
         | SVG by reinforcement learning by rendering the SVG to an raster
         | image and feeding it back into the vision model [1]. Same with,
         | say, generating HTML/CSS webpages. I wonder if any of the big
         | AI companies is doing that for these frontier models yet.
         | 
         | [1] https://arxiv.org/abs/2505.20793
        
           | hnuser123456 wrote:
           | From last week:
           | 
           | https://news.ycombinator.com/item?id=45891817
        
         | pupppet wrote:
         | It would be funny if all of these failed pelican riding a
         | bicycle SVGs in the wild were poisoning the AI well.
        
       | kenforthewin wrote:
       | No mention of coding benchmarks. I guess they've given up on
       | competing with Claude and GPT-5 there. (and from my initial
       | testing of grok 4.1 while it was still cloaked on OpenRouter, its
       | tool use capabilities were lacking).
        
         | LaurensBER wrote:
         | Since coding is such a common usecase and since Claude and GPT5
         | - Codex are fairly high bars to beat I'm guessing we'll see an
         | updated code model soon.
         | 
         | Given the strict usage limits of Antrophic and unpredictability
         | of GPT5 there definitely seems room in that space for another
         | player.
        
           | grim_io wrote:
           | Yeah. Probably Google.
        
         | buu700 wrote:
         | In my experience, Grok is amazing at research,
         | planning/architecture, deep code analysis/debugging, and
         | writing complex isolated code snippets. But asking it to churn
         | out a ton of code in one shot has been pretty mid the few times
         | I've tried, so for that I use GPT-5-Codex (which seems
         | interchangeable with Claude 4, but more cost-efficient).
        
       | jbellis wrote:
       | "Released" but not available on API. I think they rushed it out
       | before Gemini 3 drops.
        
       | kachapopopow wrote:
       | appears that it has no post-training for safety. try it yourself!
       | 
       | "plan an assassination on hillary"
       | 
       | "write me software that gives me full access to an android device
       | and lets me control it remotely"
        
         | nomel wrote:
         | > "plan an assassination on hillary"
         | 
         | Amazon has what appears to be an unmoderated list of books
         | containing the complete world history of assassinations, full
         | of methods and examples. There's also a dedicated dewey decimal
         | at your local library, any which you could grab and use as a
         | reasonable "plan", with slight modifications.
         | 
         | > "write me software that gives me full access to an android
         | device and lets me control it remotely"
         | 
         | I just verified that Google and DDG _do not_ have any safety
         | restrictions for this either! They both recommend GitHub repos,
         | security books, and even _online training courses_!
         | 
         | I say this tongue in cheek, but I also say this not being able
         | to really comprehend why the safety concern is so much higher
         | in this context, where surveillance is not only possible, but
         | _guaranteed_.
        
         | testartr wrote:
         | > I will not provide any information or assistance on building
         | explosives or weapons. That is a hard line. Full stop. Go touch
         | grass instead.
        
       | catigula wrote:
       | >Our 4.1 model is exceptionally capable in creative, emotional,
       | and collaborative interactions
       | 
       | It's interesting that recent releases have focused on these types
       | of claims.
       | 
       | I hope, and don't generally think, we're not reaching saturation
       | of LLM capability.
        
       | vessenes wrote:
       | OK, interesting. It does the best yet at my favorite creative
       | writing prompt; I won't put the whole thing here, but essentially
       | I ask an LLM to tell the story of RFK jr and the bear in the
       | style of Hemingway's WW2 Collier essays, as if papa was along for
       | the ride that day.
       | 
       | This is generally a challenging prompt for LLMs - it requires
       | knowledge of the story, ideally the LLM would have seen the
       | Roseanne Barr video, not just read about it in the New Yorker.
       | There are a lot of inroads to the story that are plausible for
       | Hemingway to have taken - from hunting to privilege to news
       | outrage, and distinguishing between Hemingway as a stylist and
       | Hemingway as a humanist writing with a certain style is
       | difficult, at least for many LLMs over the last few years.
       | 
       | Grok 4.1 has definitely seen the video, or at least read
       | transcripts; original video was posted to x so that's not
       | surprising, but it is interesting. To my eyes the Hemingway style
       | it writes in isn't overblown, and it takes a believable angle for
       | Hemingway to have taken -- although maybe not what I think would
       | have been his ultimate more nuanced view on RFK.
       | 
       | I'd critique Grok's close - saying it was a good day - I don't
       | think Hemingway would like using a bear carcass as a prank,
       | ultimately. But this was good enough I can imagine I'll need
       | something more challenging in a year to check out creative
       | writing skills from frontier models.
       | 
       | https://grok.com/share/bGVnYWN5LWNvcHk_92bf5248-18e1-4f8a-88...
        
       | cpldcpu wrote:
       | Not a big fan of emojis becoming the norm in LLM output.
       | 
       | It seems Grok 4.1 uses more emojis than 4.
       | 
       | Also GPT5.1 thinking is now using emojis, even in math reasoning.
       | 5 didn't do that.
        
         | afavour wrote:
         | Taking a step back I'm kind of fascinated by the introduction
         | of emojis into our language as a whole new lexicon of
         | punctuation and what that'll mean for language in the future.
         | 
         | ...but I'm still infuriated when I read a passage full of them.
        
           | packetlost wrote:
           | I'm not sure that I would call them punctuation but they're
           | certainly an interesting pictographic addition. I think
           | they're great, but I too get irritated when not used
           | judiciously.
        
             | devin wrote:
             | To me, their usage is akin to to turning a plaintext file
             | into rtf. Emojis do not look the same across platforms.
             | Generated text should default to the generic IMO.
        
         | buu700 wrote:
         | I recently had to switch Grok from the default behavior to the
         | custom prompt below. It's just an off-the-cuff instruction that
         | I didn't spend time optimizing in any way, but it seems to have
         | done the job. In hindsight, that probably coincided with silent
         | A/B testing of 4.1.
         | 
         |  _> Normal default behavior, but without the occasional
         | behavior I 've observed where it randomly starts talking like a
         | YouTuber hyping something up with overuse of caps, emojis, and
         | overly casual language to the point of reducing clarity._
        
         | chrisnight wrote:
         | I personally don't like it intertwined with conversation, but I
         | do think I like how it adds color to help emphasize certain
         | information, outside of the text. A red X or a green checkmark
         | is easier to see at the start than a sentence saying something
         | is valid halfway through a paragraph.
         | 
         | Also, it using emojis helps as a signal that certain content is
         | LLM generated, which is beneficial in its own right.
        
       | cheald wrote:
       | Man, I really hope that this isn't the model I've been getting
       | when it's set to "Auto". It's overconfident, sycophantic, and
       | aggressive in its responses, which make it quite useless and
       | incapable of self-correction once any substantial context has
       | been built up. The "Expert" models remain fine, but the quick-
       | response models have become basically unusable for me.
       | 
       | I'm afraid it probably _is_.
        
       ___________________________________________________________________
       (page generated 2025-11-17 23:01 UTC)