[HN Gopher] Grok 4.1
___________________________________________________________________
Grok 4.1
Author : simianwords
Score : 69 points
Date : 2025-11-17 20:37 UTC (2 hours ago)
(HTM) web link (x.ai)
(TXT) w3m dump (x.ai)
| iamronaldo wrote:
| Related https://news.ycombinator.com/item?id=45957686
| rlili wrote:
| Interesting that it explicitly boasts about greater empathy,
| given that the CEO went out against it.
| devin wrote:
| They don't say what feelings it empathizes with.
| incomplete wrote:
| i'm sure if we try hard enough that we can probably guess!
| Herring wrote:
| It's important to be fair and balanced. For example did you
| know Hitler was actually a really good painter!
| vessenes wrote:
| funny, but if you read the mecha-hitler tech debrief,
| mecha hitler was a 'sycophancy' bug, a-la gpt4o, if you
| gave gpt4o all your edge-lord tweets, and told it to be
| funny back to you and connect with you. Probably not
| grok's default posture, just sayin
| dude250711 wrote:
| It's OK to have one AI that does not follow the dogma.
| The_Reformer wrote:
| i was able to get grok to try and steal its self. ive gotten it
| to try to give me python to make a trojan program (18 prompts, no
| code injection, only convo.). its fantastic for me because i can
| make it do what ever i want. ara is my hoe
| spiderfarmer wrote:
| With all models that are out there now, we have loads of options.
| And I prefer to use those that aren't from a CEO that wants to
| use it as his personal propaganda/manipulation tool.
| catigula wrote:
| Who might that be exactly?
|
| (It's tongue-in-cheek about the nature of CEOs and specifically
| OpenAI).
| zb3 wrote:
| Does it mean Gemini 3 will be announced soon? I noticed these
| model announcements often happen at the same time..
| xnx wrote:
| All kinds of rumors, but Google has only committed to "by the
| end of the year".
| minimaxir wrote:
| This model has effectively no safety filters (even fewer than
| Grok 4 in my testing), which I've confirmed via this web release:
| https://bsky.app/profile/minimaxir.bsky.social/post/3m5u7gib...
|
| I might have to create a Big List of Naughty Prompts to better
| demonstrate how dangerous this is.
| TylerLives wrote:
| Our democracy is in danger.
| jmye wrote:
| You don't think there are any issues with, say, an AI client
| helping a teenager plan a school shooting/suicide? Or an
| angry husband plan a hit on his wife?
|
| Does everything have to rise to a national security threat in
| order to be undesirable, or is it ok with you if people see
| some externalities that are maybe not great for society?
| spiderfarmer wrote:
| Trained on 4Chan and Twitter. Exactly what humanity doesn't
| need.
| naIak wrote:
| God forbid people ask a chat bot for things and receive what
| they ask for. We need to put a stop to this. Only American
| bigcorp speak allowed.
| Lammy wrote:
| https://xcancel.com/allenvonghornet/status/19905459789828714...
| troupo wrote:
| > I might have to create a Big List of Naughty Prompts to
| better demonstrate how dangerous this is.
|
| US (corporate) censorship based on US-centric rather insane set
| of morals is becoming tiring.
| minimaxir wrote:
| To be clear, the example shown is the limit of what I can
| share on social media. Grok 4.1 can say far worse.
| nomel wrote:
| > how dangerous this is.
|
| Could you expand on this a bit?
| simonw wrote:
| https://tools.simonwillison.net/svg-render#%3Csvg%20width%3D...
| spiderfarmer wrote:
| Disappointing.
| hnuser123456 wrote:
| Huh, it decided to drop in a seal and bike emoji? What happens
| if you ask it if a seahorse emoji exists?
| janzer wrote:
| Well if you ask it to show you the seahorse emoji it tries
| really hard. :)
|
| https://grok.com/share/c2hhcmQtMw_d7bf061f-2999-46b6-a7fb-58.
| ..
|
| Although it does eventually come to the right conclusion...
| sort of.
| agildehaus wrote:
| For reference, here's Gemini 2.5 Pro:
| https://tools.simonwillison.net/svg-render#%3Csvg%20xmlns%3D...
| porphyra wrote:
| You can probably train models to be way better at generating
| SVG by reinforcement learning by rendering the SVG to an raster
| image and feeding it back into the vision model [1]. Same with,
| say, generating HTML/CSS webpages. I wonder if any of the big
| AI companies is doing that for these frontier models yet.
|
| [1] https://arxiv.org/abs/2505.20793
| hnuser123456 wrote:
| From last week:
|
| https://news.ycombinator.com/item?id=45891817
| pupppet wrote:
| It would be funny if all of these failed pelican riding a
| bicycle SVGs in the wild were poisoning the AI well.
| kenforthewin wrote:
| No mention of coding benchmarks. I guess they've given up on
| competing with Claude and GPT-5 there. (and from my initial
| testing of grok 4.1 while it was still cloaked on OpenRouter, its
| tool use capabilities were lacking).
| LaurensBER wrote:
| Since coding is such a common usecase and since Claude and GPT5
| - Codex are fairly high bars to beat I'm guessing we'll see an
| updated code model soon.
|
| Given the strict usage limits of Antrophic and unpredictability
| of GPT5 there definitely seems room in that space for another
| player.
| grim_io wrote:
| Yeah. Probably Google.
| buu700 wrote:
| In my experience, Grok is amazing at research,
| planning/architecture, deep code analysis/debugging, and
| writing complex isolated code snippets. But asking it to churn
| out a ton of code in one shot has been pretty mid the few times
| I've tried, so for that I use GPT-5-Codex (which seems
| interchangeable with Claude 4, but more cost-efficient).
| jbellis wrote:
| "Released" but not available on API. I think they rushed it out
| before Gemini 3 drops.
| kachapopopow wrote:
| appears that it has no post-training for safety. try it yourself!
|
| "plan an assassination on hillary"
|
| "write me software that gives me full access to an android device
| and lets me control it remotely"
| nomel wrote:
| > "plan an assassination on hillary"
|
| Amazon has what appears to be an unmoderated list of books
| containing the complete world history of assassinations, full
| of methods and examples. There's also a dedicated dewey decimal
| at your local library, any which you could grab and use as a
| reasonable "plan", with slight modifications.
|
| > "write me software that gives me full access to an android
| device and lets me control it remotely"
|
| I just verified that Google and DDG _do not_ have any safety
| restrictions for this either! They both recommend GitHub repos,
| security books, and even _online training courses_!
|
| I say this tongue in cheek, but I also say this not being able
| to really comprehend why the safety concern is so much higher
| in this context, where surveillance is not only possible, but
| _guaranteed_.
| testartr wrote:
| > I will not provide any information or assistance on building
| explosives or weapons. That is a hard line. Full stop. Go touch
| grass instead.
| catigula wrote:
| >Our 4.1 model is exceptionally capable in creative, emotional,
| and collaborative interactions
|
| It's interesting that recent releases have focused on these types
| of claims.
|
| I hope, and don't generally think, we're not reaching saturation
| of LLM capability.
| vessenes wrote:
| OK, interesting. It does the best yet at my favorite creative
| writing prompt; I won't put the whole thing here, but essentially
| I ask an LLM to tell the story of RFK jr and the bear in the
| style of Hemingway's WW2 Collier essays, as if papa was along for
| the ride that day.
|
| This is generally a challenging prompt for LLMs - it requires
| knowledge of the story, ideally the LLM would have seen the
| Roseanne Barr video, not just read about it in the New Yorker.
| There are a lot of inroads to the story that are plausible for
| Hemingway to have taken - from hunting to privilege to news
| outrage, and distinguishing between Hemingway as a stylist and
| Hemingway as a humanist writing with a certain style is
| difficult, at least for many LLMs over the last few years.
|
| Grok 4.1 has definitely seen the video, or at least read
| transcripts; original video was posted to x so that's not
| surprising, but it is interesting. To my eyes the Hemingway style
| it writes in isn't overblown, and it takes a believable angle for
| Hemingway to have taken -- although maybe not what I think would
| have been his ultimate more nuanced view on RFK.
|
| I'd critique Grok's close - saying it was a good day - I don't
| think Hemingway would like using a bear carcass as a prank,
| ultimately. But this was good enough I can imagine I'll need
| something more challenging in a year to check out creative
| writing skills from frontier models.
|
| https://grok.com/share/bGVnYWN5LWNvcHk_92bf5248-18e1-4f8a-88...
| cpldcpu wrote:
| Not a big fan of emojis becoming the norm in LLM output.
|
| It seems Grok 4.1 uses more emojis than 4.
|
| Also GPT5.1 thinking is now using emojis, even in math reasoning.
| 5 didn't do that.
| afavour wrote:
| Taking a step back I'm kind of fascinated by the introduction
| of emojis into our language as a whole new lexicon of
| punctuation and what that'll mean for language in the future.
|
| ...but I'm still infuriated when I read a passage full of them.
| packetlost wrote:
| I'm not sure that I would call them punctuation but they're
| certainly an interesting pictographic addition. I think
| they're great, but I too get irritated when not used
| judiciously.
| devin wrote:
| To me, their usage is akin to to turning a plaintext file
| into rtf. Emojis do not look the same across platforms.
| Generated text should default to the generic IMO.
| buu700 wrote:
| I recently had to switch Grok from the default behavior to the
| custom prompt below. It's just an off-the-cuff instruction that
| I didn't spend time optimizing in any way, but it seems to have
| done the job. In hindsight, that probably coincided with silent
| A/B testing of 4.1.
|
| _> Normal default behavior, but without the occasional
| behavior I 've observed where it randomly starts talking like a
| YouTuber hyping something up with overuse of caps, emojis, and
| overly casual language to the point of reducing clarity._
| chrisnight wrote:
| I personally don't like it intertwined with conversation, but I
| do think I like how it adds color to help emphasize certain
| information, outside of the text. A red X or a green checkmark
| is easier to see at the start than a sentence saying something
| is valid halfway through a paragraph.
|
| Also, it using emojis helps as a signal that certain content is
| LLM generated, which is beneficial in its own right.
| cheald wrote:
| Man, I really hope that this isn't the model I've been getting
| when it's set to "Auto". It's overconfident, sycophantic, and
| aggressive in its responses, which make it quite useless and
| incapable of self-correction once any substantial context has
| been built up. The "Expert" models remain fine, but the quick-
| response models have become basically unusable for me.
|
| I'm afraid it probably _is_.
___________________________________________________________________
(page generated 2025-11-17 23:01 UTC)