[HN Gopher] Player of Games
       ___________________________________________________________________
        
       Player of Games
        
       Author : vatueil
       Score  : 341 points
       Date   : 2021-12-08 05:52 UTC (17 hours ago)
        
 (HTM) web link (arxiv.org)
 (TXT) w3m dump (arxiv.org)
        
       | crhutchins wrote:
       | I'll try to look into a brighter light into this one.
        
       | SuoDuanDao wrote:
       | I didn't even know about the book until I read the comments here,
       | I thought it was a reference to the Grimes song. Funny
       | coincidence the song and the engine would appear so close in time
       | to one another.
        
         | Severian wrote:
         | The Grimes song is a reference to the book too. She also has
         | Marain subtitles in her video for "Idoru", which is the
         | language used in The Culture. Weird mix of two author's (Idoru
         | being William Gibson) works to be sure.
        
       | mudlus wrote:
       | Yawn, show me a computer that game make fun games
        
         | TaupeRanger wrote:
         | You're getting downvotes but honestly I agree. Who cares about
         | board games? We should've moved on from this once we "solved"
         | chess and Go. There are more important things and it's not
         | remotely surprising that a computer can beat a human when
         | there's a simple, abstract optimization problem to throw
         | computing power at. Make it creative...now that's a challenge
         | worthy of the top AI talent.
        
           | newswasboring wrote:
           | I agree. I have always wondered if I can feed GPT-3 a bunch
           | of rule books and ask it to generate game rules.
        
           | kadoban wrote:
           | You haven't seen AlphaGo play Go then, it plays creatively as
           | hell at points.
        
       | [deleted]
        
       | [deleted]
        
       | skinner_ wrote:
       | It would be awesome to have two interacting communities: AI
       | experts building open source general game playing engines, and
       | gaming fans writing pluggable rule specifications and UIs for
       | popular games.
       | 
       | A bit of googling shows that there is a General Game Playing AI
       | community with their own Game Description Language. I never
       | really encountered them before, and the DeepMind paper does not
       | cite them, either.
        
         | dpflug wrote:
         | Last I looked, the GGP community is focused on perfect
         | information games currently. I had the same thought, though.
        
       | sfkgtbor wrote:
       | I really like seeing references to the Culture series when naming
       | things:
       | 
       | https://en.m.wikipedia.org/wiki/The_Player_of_Games
        
         | doctor_eval wrote:
         | I suppose it's better than "Use of Weapons".
        
           | OneTimePetes wrote:
           | Why not have a seat, take that chair over there.
        
             | _0ffh wrote:
             | One of the best, and executed to perfection! You can sort-
             | of-see the point coming for a long, long time in the book,
             | as he gradually builds the suspicion by dropping the
             | occasional hint here and there, but it's always so that it
             | must remain a highly uncertain speculation until he drops
             | the reveal. Just the right balance between "How should I
             | have suspected that?" and "Those hints were too much on the
             | nose!".
        
               | OneTimePetes wrote:
               | Its such a crime - of war and all else, its like a
               | blindspot of imagination. That a man would do such a
               | thing - to what is essentially family, as tactics.. the
               | horror..
        
         | dane-pgp wrote:
         | I think it is also a reference to "PogChamp", although it's
         | disappointing that PoG apparently wasn't evaluated against the
         | Arcade Learning Environment (ALE) corpus of Atari 2600 games.
        
           | abledon wrote:
           | much more refined to think a spam of "POG!" stands for Player
           | of Games when reading twitch chat
        
         | hoseja wrote:
         | Kinda ironic since in the novel, a human player is better than
         | the strong AI (albeit a little inexplicably).
        
           | 7thaccount wrote:
           | I thought the protagonist wasn't nearly as talented as the
           | culture AIs (even the ones that are not all that powerful)?
        
             | thom wrote:
             | Is that clear from the text? Gurgeh supposedly perceives
             | the result of the last game before the AIs so we're led to
             | believe he's seeing deeper. Obviously he could have been
             | wrong and still won. The AIs lied to and manipulated him
             | the entire time so it's hard to know, but it would seem a
             | very odd weakness for an AI to have. I think Banks pretty
             | quickly recanted on the subject of the Culture's
             | 'referrers' but I don't think he plays a full Mind, so it's
             | not a clear cut conversation.
        
               | joshuamorton wrote:
               | My recollection is that by the end of the novel its clear
               | that Gurgeh was never competitive with the ship, although
               | he might have been competitive with his security drone
               | (although even that isn't clear, since <spoilers> imply
               | that the security drone is a better game player than it
               | pretends to be).
               | 
               | To me it felt like the whole point of the novel was that
               | Gurgeh was a piece in an even larger game and he didn't
               | even realize it. So the idea that the people playing the
               | "bigger" game couldn't compete in the smaller game seems
               | silly, and I think they mention that they used Gurgeh
               | instead of an AI to make it appear fair to the
               | inhabitants of the planet.
        
               | hesperiidae wrote:
               | I agree with you that Gurgeh was just a piece getting
               | manipulated and that that was the point, but Gurgeh was
               | still the best piece _that they could use_ for the job.
               | 
               | The Culture is (in this story) pretty much only bound by
               | their own constraints. They chose Gurgeh for the role,
               | since he had enough skill and talent to actually be able
               | to accomplish the Culture's (or the SC's, _winkwink_ )
               | objectives without having the whole thing being taken
               | over by an AI.
               | 
               | The Culture worked very much like the PoG this thread is
               | about: it minimised potential loss and considered the
               | constraints it had to get the best possible outcome.
               | 
               | The Culture is mostly constrained by only ethical rules,
               | which, admittedly, can get flexible, especially with
               | regards to the SC. The practical restrictions, like it
               | being easier to send one capable human than to conquer a
               | small galaxy, are in my mind lesser in comparison.
               | 
               | As such, I think they got the most out of the operation,
               | just by being confident in their assessment of a single
               | human who played games good. And there's absolutely no
               | reason to believe that the overminds that guide the
               | Culture can't model human behaviour down to the smallest
               | variable, especially considering how augmented humans are
               | in the Culture.
               | 
               | I'm also 100% onboard the idea that all the drones could
               | outplay Gurgeh in a blink in any game, intuition be
               | damned.
        
               | sdenton4 wrote:
               | Yeah, I thought it was clear from the beginning of the
               | book that no humans were even remotely competitive with
               | any AI (including the main character) but that human game
               | players were sort of an aesthetic throwback, like dog-
               | racing in an era of F1 cars.
        
               | 7thaccount wrote:
               | This was my understanding as well, but I might have read
               | into it. The culture minds are in freaking hyperspace to
               | get around lightspeed limitations on computations. He for
               | sure can't beat that, but he could beat someone on
               | another planet at their own game that he literally just
               | learned in the year it took to get there. A game that
               | permeates every aspect of their civilization.
               | 
               | I do assume his drone could beat him as well, but I'm not
               | sure.
        
             | hoseja wrote:
             | I don't think a full Culture Mind is present but he
             | outstrips his spacecraft's ability to help him with
             | preparation in later stages of the competition. I clearly
             | remember this.
        
               | macmac wrote:
               | At least that is what the ship (SC) wants him to think.
        
               | WJW wrote:
               | Indeed. (spoiler following) The plot basically revolves
               | around SC manipulating both Gurgeh and the Empire of Azad
               | in an ever bigger and complex game than the one in the
               | book. Given how Banks describes the Minds in other books
               | it would be extremely curious if they wouldn't crush any
               | biological player in any normal game the same way chess
               | computers crush humans these days. But, it is possible
               | that a more limited mind like the security drone could be
               | outstripped by Gurgeh. In one of the other books they do
               | mention that "smaller" machines like environmental suits
               | and small drones get more limited minds than full
               | starships as it would be cruel to put a fully capable
               | Mind in such a limited body.
        
               | stavros wrote:
               | How did I miss this plot point? It's been a while, but I
               | remember focusing on the game Gurgeh played. Maybe I just
               | don't remember it now.
        
               | WJW wrote:
               | The last page of the book gives it away: (MASSIVE SPOILER
               | OBV) The security drone who came with Gurgeh to Azad was
               | the same drone he meets during the introduction chapters
               | who was "rejected" from SC and offers to let him cheat
               | (though it was wearing a disguise at the time). Then,
               | after he cheats he basically gets blackmailed into going
               | to Azad and conveniently this "non-SC" drone comes with
               | him in a very "non-SC" ship that claims to have its
               | weapons removed but doesn't. At some crucial points the
               | security drone influences Gurgeh to play the best he can,
               | such as when he takes him on a tour of the slums and the
               | Culture-educated Gurgeh gets so furious at the
               | mistreatment he witnesses that he absolutely crushes his
               | opponent in the next match.
               | 
               | They mention in one of the final chapters that the minds
               | wanted the Azad empire to become a better place since it
               | was really shitty to its citizens. However, they couldn't
               | just invade and impose laws because they're the Culture,
               | and the Azad empire kept claiming moral superiority
               | because they had this one thing (the Game) that they
               | thought the Culture couldn't match. The Minds knew Gurgeh
               | was talented enough to get far enough in the tournament
               | that the Azad Empire would be seriously shaken, because
               | if this single foreigner can beat so many of the best and
               | brightest in the Empire at the thing it claims to do best
               | then what could the entire Culture do? This turns out to
               | have been correct, at the end of the book the Azad empire
               | starts to collapse because they no longer trust their
               | leadership, who have been proven to be incompetent at the
               | very thing they claim to do best. Beaten by a human btw,
               | not even by one of the god-machines that the Culture also
               | has. Having predicted that this would happen, the Minds
               | set out to manipulate Gurgeh into going to Azad to play
               | the Game and by doing so bring about regime change. The
               | Minds and/or SC were playing a much higher level game
               | than Gurgeh all along, he was merely one of the pieces
               | they used to play.
        
               | stavros wrote:
               | Ahh, thank you! Now that you recount it, it all comes
               | back to me. I should read more Banks, he's a fantastic
               | writer.
        
               | 7thaccount wrote:
               | Yes, at the end you start to question just who the player
               | of games actually was.
        
           | pharmakom wrote:
           | No he is not, but AIs are not allowed in the competition the
           | story centers around.
        
             | hoseja wrote:
             | Near the end of the competition, as he is deep in his
             | analysis, the light craft AI gives up on helping him since
             | it gets overwhelmed. Granted it's not a full Culture Mind
             | (kinda hazy, been a while) but still a point for the
             | meatbag.
        
               | pharmakom wrote:
               | I think the main character can be so strong at the game
               | by the end because of his immersion in Empire culture.
               | The ships AI would likely be at least as strong with the
               | same experiences. Plus, as you mention, the ships AI is
               | not smartest AI around.
        
               | arlort wrote:
               | I always interpreted the end reveal as showing that
               | control was highly confident of both the outcome of the
               | game and of how Gurgeh got to that outcome.
               | 
               | It's been a while but I am pretty sure that the ship lied
               | when saying that it got overwhelmed and did so only
               | because it was confident he was on the right path but
               | needed to get there in a specific way which wouldn't have
               | worked quite the same if the ship intervened
        
               | hesperiidae wrote:
               | Yeah, he wouldn't have reached such a good solution with
               | help, and that was also originally taken into account by
               | the Culture when they sent him out in the first place,
               | since they knew him _that_ thoroughly.
        
               | bduerst wrote:
               | Yep, basically the nebulous, unknown minds of Control
               | predicted the main character would win, and set up as
               | many conditions as possible to push him to do so.
               | Including bluffing about help from the AI.
               | 
               | It was part of an even bigger game but I'm not going to
               | get into spoilers.
        
         | CobrastanJorji wrote:
         | Allusions are fun and all, but I disagree. These are important
         | problems that a lot of people have put their whole careers into
         | researching. Silly names like these lack gravitas.
        
           | 0_gravitas wrote:
           | indeed
        
           | ZeroGravitas wrote:
           | Very little Gravitas Indeed.
        
             | 0_gravitas wrote:
             | Ah so its __you__ that took that one
        
           | moritonal wrote:
           | Sorry, to explain the joke. The ships name themselves, and
           | when they pick jokey names they're often mocked by the humans
           | (which are in every way essentially ants to the spaceships)
           | for not having enough gravitas. So the ships start naming
           | themeselves things like the "Death-ray 9000 super-killer
           | deluxe", to essentially take the piss.
           | 
           | Funnily enough you can see the exact same effect in principal
           | game-engineers or computer-hacking.
        
             | robbie-c wrote:
             | I believe the user you are replying to was also joking,
             | given that many of Banks' ship names reference the g-word
             | 
             | Edit: if not that's even more amusing
        
               | marvin wrote:
               | I'll make a minor contribution to the discussion by
               | mentioning the Culture ship normally referred to as the
               | Mistake Not..., which is shorthand for
               | 
               | "Mistake Not My Current State Of Joshing Gentle
               | Peevishness For The Awesome And Terrible Majesty Of The
               | Towering Seas Of Ire That Are Themselves The Milquetoast
               | Shallows Fringing My Vast Oceans Of Wrath".
               | 
               | Unsure if this name is also a sarcastic stab that the
               | lack of gravitas in ships' names, but regardless it's
               | very sad that Banks died young :(
        
               | hesperiidae wrote:
               | Yeah, it's making fun of the human desire for gravitas
               | when it comes to ship names, since it's just
               | exponentially more and more over the top.
        
           | gremloni wrote:
           | If anything the caliber and lore of the series gives the
           | project an incredible amount of gravitas. Plus the scheme is
           | just plain beautiful in my opinion.
        
             | lacker wrote:
             | You may find this Iain Banks interview enjoyable. TLDR
             | search for "gravitas" ;-)
             | 
             | https://www.theguardian.com/books/2000/sep/11/iainbanks-
             | scie...
        
               | macmac wrote:
               | Brilliant. "Absolutely No You-No-What" is fantastic.
        
               | stavros wrote:
               | Nothing can top "Ultimate Ship the Second".
        
               | xenophonf wrote:
               | My favorite is _Eschatologist_ (temporary name) lolololol
        
           | sjg1729 wrote:
           | Always sad to see these projects suffer from A Shortfall of
           | Gravitas
        
             | auggierose wrote:
             | I see what you did there :-)
        
             | NoGravitas wrote:
             | Gravitas? What Gravitas?
        
         | Borrible wrote:
         | Banks should have named one of Culture's General System
         | Vehicles 'Don't be Evil'.
         | 
         | https://theculture.fandom.com/wiki/List_of_spacecraft
        
       | wiz21c wrote:
       | Couldn't resist :
       | 
       | https://www.youtube.com/watch?v=-1F7vaNP9w0
        
       | sdenton4 wrote:
       | This is clearly part of DeepMind's long-game plan to achieve
       | world domination through board game mastery. Naming the new
       | algorithm after the book is a real tip of their hand...
       | 
       | https://en.wikipedia.org/wiki/The_Player_of_Games
        
         | chrisweekly wrote:
         | PSA: The "Culture" novels by Iain M Banks are fantastic and can
         | be read in any order. "Player of Games" was the 1st one I read
         | and still probably my favorite.
        
           | bewaretheirs wrote:
           | I keep hearing recommendations for the Culture books so I
           | tried reading it recently and it just didn't work for me -- I
           | gave up on it halfway through, which is rare for me.
        
             | pault wrote:
             | Which one? They each have a unique feel and setting.
        
             | Jordanpomeroy wrote:
             | They are a slow burn, but the ends always justify the means
             | with those novels. If you really did make it 1/2 way, I'd
             | encourage you to go back and finish reserve judgment.
        
           | bduerst wrote:
           | _Player of Games_ is the second book, and the one I recommend
           | people start _The Culture_ series with.
           | 
           | The first book _Consider Phlebas_ isn 't bad, but it isn't as
           | well developed as the rest of the series IMO.
        
             | hesperiidae wrote:
             | It's a great starting point, since not only is the story
             | both fun and interesting, but it also shows what the
             | Culture's values and methods are in a very satisfying way
             | by juxtaposing them against the Empire through the
             | tournaments of the latter's own game.
        
           | kmtrowbr wrote:
           | Yes! I love this one. It's my favorite too.
        
         | [deleted]
        
         | 7thaccount wrote:
         | Pretty amazing book. I wish I could play a board game like that
         | as well.
        
           | automatic6131 wrote:
           | I always imagine the board game as essentially being SM's
           | Civilisation but really, really good in an indescribable way
           | - with some card games inbetween.
        
             | 0_gravitas wrote:
             | I believe Banks himself said that he used to play Civ and
             | took some inspiration from it
        
               | bduerst wrote:
               | Definitely for his later books, but _Player of Games_
               | came out three years before the first Sid Meier 's Civ
               | game.
               | 
               | What cool is that Paradox's _Stellaris_ , a civ-in-space
               | game, definitely takes pages from Ian's Culture series.
        
           | stavros wrote:
           | I second this, it was excellent. I've only read a few Banks
           | books, but this was my favorite.
        
             | arvinsim wrote:
             | I started with Consider Phlebas but stopped because it
             | seems too slow for me.
             | 
             | Does it get better in the later chapters?
        
               | sidibe wrote:
               | Use of Weapons and Consider Phlebas are the worst of the
               | series IMO. I powered through Consider Phloebus just
               | because I knew people loved the series, but there's
               | really no reason to start there
        
               | rishav_sharan wrote:
               | Sacrilege! Taste is subjective, but Use of Weapons, imo,
               | is Banks best work. I personally consider it one of the
               | best SciFi, period. It's been years I last read it, but
               | that ending still gives me shivers whenever I think of it
        
               | arethuza wrote:
               | Personally, I think _Use of Weapons_ is by far the best
               | of the Culture series - although I admit it took a few
               | readings for me to get to that view...
        
               | User23 wrote:
               | He does a really good job exploring the theme "what is a
               | weapon?"
        
               | hesperiidae wrote:
               | Yes! And also, "why is a weapon?", "how is a weapon?" and
               | "what isn't a weapon?"
               | 
               | Really, Culture is a great setting for discussing a lot
               | of interesting philosophical questions and topics.
        
               | NoGravitas wrote:
               | Use of Weapons is generally considered one of the best,
               | but it has a complex narrative structure that makes it a
               | harder read, and it's probably not a good place to start.
        
               | vermilingua wrote:
               | It does, but IMO it's probably worth reading The Player
               | of Games or Use of Weapons before it anyway. With the
               | exception of perhaps Surface Detail, none of the Culture
               | books rely on any others. Consider Phlebas gives a good
               | view of The Culture from "outside" (the perspective of
               | the Idrians) but is quite slow.
        
               | DylanSp wrote:
               | I mean, Player of Games has a pretty slow start too. I
               | love that book, but the initial pacing is IMO its biggest
               | flaw.
               | 
               | I know Use of Weapons doesn't depend on any of the other
               | books for its plot, but is it a decent intro to the
               | setting? If it is, that's where I'd recommend starting.
        
               | avemg wrote:
               | I've read (in order) Consider Phlebas, Player of Games,
               | Use of Weapons, and Excession thus far. Use of Weapons
               | was the toughest one for me to get through so far. I
               | started it and stopped it a few times over several years
               | and just couldn't get past the halfway point. I
               | eventually got over the hump with it and devour the last
               | half of the book over a couple of days (which is fast for
               | me). So for my money, Use of Weapons is a bad starting
               | point.
               | 
               | My favorite by far is Excession but I don't know that I'd
               | start there. I think the payoff of getting a story from
               | the perspective of the Minds is better appreciated after
               | you've heard about them and their capabilities from a
               | distance in the preceding books.
               | 
               | My pick would be to start with Player of Games. That's
               | the one that was a page turner for me nearly from the
               | jump.
        
               | hesperiidae wrote:
               | Yeah, it does take its time to build up, but in my
               | opinion that just lays a stronger foundation for the
               | latter half(-ish) of the book.
               | 
               | Use of Weapons is more than a decent intro, but I'd still
               | personally recommend The Player of Games since it isn't
               | as deep and heavy in comparison, and the narrative
               | structure is simpler.
               | 
               | Of course, YMMV, but I started with TPoG, and reading UoW
               | right after it was absolutely fantastic. I guess my
               | biggest concern with recommending UoW as the starter
               | would be that it might diminish TPoG, which I'm fond of,
               | but I don't know if it actually would, since they're
               | connected pretty much only bound by the setting.
        
               | smiley1437 wrote:
               | If you find Consider Phlebas too slow, try Excession
               | before giving up on Banks.
        
               | speed_spread wrote:
               | Eeh, Excession is very good but a still bit hermetic for
               | an introduction to the Culture. It's the only other book
               | I wouldn't recommend as a first along with Consider
               | Phlebas.
        
               | apetersonBFI wrote:
               | Consider Pheblas wasn't the best in my opinion. Player of
               | Games is my favorite, but some of the other later ones
               | are good.
               | 
               | I enjoy the settings and the concept of the Culture more
               | than the plotlines.
        
               | hesperiidae wrote:
               | The Player of Games is my favourite exactly for the same
               | reason: it explores the Culture in contrast to the
               | Empire, and even the drama is just an expression of the
               | clash between the two structures.
        
               | dgritsko wrote:
               | I started with Consider Phlebas because I wanted to see
               | why people raved about the Culture series. I found it
               | kind of tedious and slow, and although I finished it, I
               | wondered what all the hype was about. I'm thankful that I
               | picked up the second book though (Player of Games) -
               | because I couldn't put it down; it was fantastic. I've
               | stuck with the series since then (Excession was another
               | highlight). I'd would like to revisit Consider Phlebas at
               | some point, I think I might enjoy it more now that I have
               | more context for the story.
        
               | gpderetta wrote:
               | I enjoyed Consider Phlebas a lot, but it is very
               | different from the rest of the series.
        
               | baq wrote:
               | Consider Phlebas is easily the worst of the series. My
               | top 3 in no particular order are Use of Weapons, Player
               | of Games and Excession.
        
               | DylanSp wrote:
               | Curious that you put Consider Phlebas behind Matter (my
               | least favorite, by far). My favorite is probably Look to
               | Windward, closely followed by Player of Games and Use of
               | Weapons.
        
               | gman83 wrote:
               | I couldn't get through the cannibal part of Consider
               | Phlebas, was soo weird.
        
               | DylanSp wrote:
               | Yeah, that was just...out-of-nowhere gruesomeness.
        
               | wishinghand wrote:
               | I guess I just took it as another possibility in a
               | society of nearly infinite ones. I did use material from
               | that encounter in running a horror RPG, so in a way I'm
               | kind of thankful for it.
        
               | wiredfool wrote:
               | There's one scene or so in each one of his books that's
               | just too much for me. I just don't need to donate
               | brainspace to that sort of thing. (Use of Weapons has
               | one, Song of Stone too.)
               | 
               | I like 80% of his work, 15% is a pointless depressing
               | slog, and the other 5% is just too much for me.
        
               | bewaretheirs wrote:
               | That confirms my decision to abandon reading the series
               | after I hit a spot like that in Player.
        
         | sillysaurusx wrote:
         | The abbreviation is PoG too. I bet that was totally on purpose.
         | At least one person in Brain is a dota player, so you better
         | believe they watch twitch.
         | 
         | Funny that most of the comments are about the name. What an
         | excellent choice.
        
         | [deleted]
        
         | WithinReason wrote:
         | "In 2015, two SpaceX autonomous spaceport drone ships--Just
         | Read the Instructions and Of Course I Still Love You--were
         | named after ships in the book, as a posthumous tribute to Banks
         | by Elon Musk"
        
           | omnicognate wrote:
           | Shame they didn't go with Pure Big Mad Boat Man.
        
         | 6510 wrote:
         | The end game is pinball and we are the balls.
        
           | zeristor wrote:
           | We are the pins.
        
       | pixelpoet wrote:
       | Anyone else surprised to see that Demis Hassabis didn't have a
       | hand in this research? Given his background as a player of many
       | games, and involvement in a lot of their research.
        
       | [deleted]
        
       | loxias wrote:
       | Psh, wake me when it can play Mao. ;)
        
       | WilliamDampier wrote:
       | so this is what Grimes latest song is about?
        
         | junon wrote:
         | Yeah wtf, was my first thought. This is mind blowing if true.
        
           | junon wrote:
           | Actually she probably got the name from the sci-fi novel this
           | is named after:
           | 
           | https://en.m.wikipedia.org/wiki/The_Player_of_Games
        
         | 323 wrote:
         | > All the lyrical evidence that Grimes' new song 'Player of
         | Games' is about ex Elon Musk
         | 
         | > Grimes seemingly makes multiple, thinly veiled references to
         | Musk in the song
         | 
         | https://www.independent.co.uk/arts-entertainment/music/news/...
        
           | cwkoss wrote:
           | SpaceX's landing pad barges are also named after Culture
           | series starships
        
       | hervature wrote:
       | I think this is a good step forward that generalizes an algorithm
       | to play both perfect and imperfect information games. However,
       | table 9 shows (I believe it shows, it is not the most intuitive
       | form), that other AIs (Deepstack, ReBeL, and Supremus) eat its
       | lunch at poker. It also performs worse than AlphaZero at perfect
       | information games. So, while a nice generalizing framework,
       | probably will not be what you use in practice.
        
       | [deleted]
        
       | wly_cdgr wrote:
       | The future is so depressing
        
         | wetpaws wrote:
         | Fun fact: The consensus between professional go and chess
         | players is that all new AI systems (alphago, etc) have really
         | revitalised the game and introduced incredible amount of new
         | strategies and depth.
        
           | jart wrote:
           | Sad fact: Lee Sedol retired after AlphaGo defeated him.
        
             | jm547ster wrote:
             | 3 years after...
        
               | newswasboring wrote:
               | yeah but still because of it[1]
               | 
               | [1] https://www.theverge.com/2019/11/27/20985260/ai-go-
               | alphago-l...
        
               | dane-pgp wrote:
               | I don't want to pull back the curtain too much, but
               | surely DeepMind foresaw the possibility of AlphaGo
               | winning and then Lee Sedol losing confidence or interest
               | in the game, which would generate a load of bad publicity
               | for them.
               | 
               | So it would make sense for DeepMind's contract with him
               | to contain a clause requiring him to continue playing go
               | professionally for a few years (but not necessarily put
               | much effort into it), as well as the standard non-
               | disparagement clauses.
               | 
               | In fact, I wouldn't be surprised if AlphaGo was
               | programmed to throw the forth game after securing the win
               | with the first three of the five games. That gives Lee
               | some bragging rights, and makes for a more hopeful story
               | than "Computer stomps likeable human".
        
               | jart wrote:
               | Here's what Lee Sedol said when he retired:
               | 
               | > With the debut of AI in Go games, I've realized that
               | I'm not at the top even if I become the number one
               | through frantic efforts.
               | https://en.yna.co.kr/view/AEN20191127004800315
               | 
               | He'd been playing Go professionally for 24 years. I never
               | said he ragequit. He's too great a man to do something
               | like that. Lee instead apologized for his losses, stating
               | "I misjudged the capabilities of AlphaGo and felt
               | powerless" while emphasizing that the defeat was his own
               | and "not a defeat of mankind". I imagine being the Hector
               | of humanity is quite a burden to bear. His professional
               | ratings then took a dive for a few years
               | https://www.goratings.org/en/history/ before he announced
               | his retirement. To this day he remains the only human
               | being who's ever won a single game against AlphaGo.
        
             | visarga wrote:
             | Caching out at the height of his fame.
        
           | loxias wrote:
           | I wish alphago was more "democratized" -- that is to say, I
           | have many questions and experiments I'd love to run on it (a
           | friend of mine and I have frequently pondered Go played in
           | various different topological spaces, and I'd love to see an
           | AI's result, for example).
        
             | kadoban wrote:
             | Look into Katago. It's an open source AI in the same
             | general style as AlphaGo, with an empasis on training
             | speed. On 9x9 you can get up to superhuman really quickly
             | on just a decent home machine (I think hours/days, can't
             | remember exactly and it's probably improved since I
             | looked).
        
               | elefantastisch wrote:
               | You can also just download pre-trained models. Get those
               | set up and then install Sabaki
               | (https://sabaki.yichuanshen.de/) and connect it to your
               | KataGo... instant (ok, a few hours probably if it's your
               | first time setting it up) superhuman Go AI. There's even
               | an npm package you can use to process SGF files and
               | automatically score moves as good/questionable/bad +
               | generate variations that were better choices:
               | https://github.com/9beach/analyze-
               | sgf/blob/master/README.en-...
               | 
               | (Edit: Misread what the other poster was trying to do,
               | but I'm leaving this here as a reference for anyone else
               | who just wants to use KataGo on their own machine on
               | their own Go games.)
        
               | loxias wrote:
               | Awesome! Thanks! Checking it out now.
        
             | franknstein wrote:
             | Fun idea. Did you reach any interesting conclusions?
        
           | wly_cdgr wrote:
           | Yeah, whatever. As someone who grew up playing chess and is
           | almost certainly much better at it than you, this future
           | sucks
        
             | wsc981 wrote:
             | I don't understand why this is so depressing? You can still
             | play against humans though. It's probably more fun anyways
             | than playing versus a computer as in most games, isn't it?
        
               | logicchains wrote:
               | Like in first-person shooters; it's no fun to play
               | against bots with perfect aim and superhuman reflexes.
        
               | WJW wrote:
               | Then don't, there are plenty of humans available to play
               | chess (or first person shooters) with.
        
               | krageon wrote:
               | But it is fun to play against well-tuned bots. That's one
               | of the ways you can improve.
        
             | AlexAndScripts wrote:
             | Why?
        
               | Kaibeezy wrote:
               | Because the only game left will be thrones?
        
               | JanneVee wrote:
               | I don't know how go changed. But as for chess the
               | tournament play at the master level have insane deep
               | opening preparation done before with computers. They play
               | preparation game where they try to guess what lines the
               | opponent checked and memorized before the games. They
               | aren't actually playing until their computer backed
               | preparation ends more than the few moves that they have
               | fed in to come up with something different. Both
               | spectators and players kind of find this a little bit
               | boring.
               | 
               | I do acknowledge that this isn't a new phenomena Fischer
               | complained about this before the computer engine era and
               | came up with a chess variant to nullify deep opening
               | prep!
        
               | Kelamir wrote:
               | The change in Go is that professionals now play more AI-
               | like stuff, basically the same opening moves, and that's
               | fairly boring to watch. It's the best sequences of moves
               | we have so far, but also everyone knows them and isn't
               | interested.
               | 
               | Another change is that territory is more valued than
               | influence now, which too makes games less fun to watch,
               | at least in my experience. To my knowledge, Shibano
               | Toramaru, professional Go player, used to play highly
               | focused on influence, and his games were very interesting
               | to watch; it was just spectacular. But after AlphaGo came
               | he converted to focus on territory like everyone else,
               | only occasionally letting his beast out. But I watched
               | only a few videos on him so don't take my word; it's just
               | my impression.
        
               | _0ffh wrote:
               | Yeah, I guess half the opening is hoping to kick the
               | opponent out of his preparations while staying within
               | your own. Or maybe all.
        
               | zem wrote:
               | climate change, no doubt.
        
       | cab404 wrote:
       | SCP-like name for SCP-like neural network.
       | 
       | "SCP-29123 Player Of Games"
        
       | ArtWomb wrote:
       | This seems like a significant milestone in AI. I mean what can't
       | an agent with mastery of "guided search, learning, and game-
       | theoretic reasoning" accomplish?
        
         | ausbah wrote:
         | modeling every task as a game seems like a big hurdle, or even
         | just getting a working "environment"
        
       | RivieraKid wrote:
       | Wow, it can beat a good poker bot, that is impressive.
        
       | BeenChilling wrote:
       | I want to see deepmind make a bot to play team based first person
       | shooters like csgo and rainbow6 siege, to stack up five of them
       | against a team of professional players.
        
         | mensetmanusman wrote:
         | They probably won't for publicity reasons.
        
         | ausbah wrote:
         | that's what OpenAI did a couple yewrs ago with Dota 2
         | 
         | https://openai.com/five/
        
         | fho wrote:
         | Honestly that probably won't be too interesting as (a) one AI
         | could perfectly control several agents (ie perfect coordination
         | of global strategies) and (b) an AI has low to no reaction
         | times and perfect aim (aimbots already have that) so I would
         | expect that would quickly result in a slaughterfest.
        
           | LudwigNagasena wrote:
           | (a) make them independent (b) add 100-200ms delay
        
           | arethuza wrote:
           | _"...such consummate skill, such ability, such adaptability,
           | such numbing ruthlessness, such a use of weapons when
           | anything could become weapon... "_
        
           | gverrilla wrote:
           | Same applies to dota2, and it was very interesting what they
           | did there. But yeah first they would need to simulate how
           | human players react and aim, or it would be impossible to
           | play against.
        
           | ausbah wrote:
           | IIRC multi-agent domains are in their own category
           | specifically because a single agent posing as "multiple
           | agents" usually can't solve such environments, you need
           | multiple agents with varying degrees of dependence
        
           | arlort wrote:
           | What would be interesting would be 5 independent AIs (even
           | just different instances of the same AI of course) using the
           | same interface as human players, so the same controls and the
           | same video output
           | 
           | I am pretty sure aimbots access the internals of the game
           | rather than reading the video output to identify the
           | silhouette of the enemy.
        
       | bkartal wrote:
       | Impressive work! Most authors, if not all, are from DeepMind
       | Edmonton office.
        
         | cmauniada wrote:
         | I didn't even know that they had an office in Edmonton...
        
           | [deleted]
        
           | bkartal wrote:
           | Edmonton is one of the best places for RL research &
           | ecosystem, both DeepMind and University of Alberta are there.
        
       | tsbinz wrote:
       | Comparing against Stockfish 8 in a paper released today and
       | labeling it as "Stockfish" is bordering on being dishonest. The
       | current stockfish version (14) would make AlphaZero look bad, so
       | they don't include it ...
        
         | [deleted]
        
         | moondistance wrote:
         | The abstract clearly states that the best chess and Go bots are
         | not beaten: "Player of Games reaches strong performance in
         | chess and Go, beats the strongest openly available agent in
         | heads-up no-limit Texas hold'em poker (Slumbot)..."
        
         | ShamelessC wrote:
         | The first mention says "Stockfish 8, level 20" in the paper.
         | This isn't a blog post that you can skim, you need to read the
         | whole thing before critiquing.
        
           | tsbinz wrote:
           | I obviously read it, otherwise I wouldn't have known which
           | version they are using. They are banking on others, that do
           | just skim the figures and tables, not noticing their usage of
           | outdated baselines.
        
           | karpierz wrote:
           | That's actually the second mention, the first is when they
           | introduce the games in section 4:
           | 
           | > Today, computer- playing programs remain consistently
           | super-human, and one of the strongest and most widely-used
           | programs is Stockfish.
           | 
           | They also go back to referring to it as Stockfish for the
           | rest of the paper.
           | 
           | An analogous situation in my mind would be if AMD released a
           | new CPU and benchmarked it against an Intel CPU, only
           | mentioning once, somewhere in the middle of the paper, that
           | it was a Pentium 4.
        
             | ShamelessC wrote:
             | > Today, computer- playing programs remain consistently
             | super-human, and one of the strongest and most widely-used
             | programs is Stockfish.
             | 
             | This is just a general effort to describe the present state
             | of things. When they explicitly describe their evaluation
             | process, they are sure to use the version number. They then
             | _immediately_ drop the version number in subsequent usage
             | which is culturally standard in research papers so they
             | don't concern themselves with minute details of every
             | single thing they find themselves redescribing. Believe me,
             | you don't want to read the verbose version of this
             | paragraph.
             | 
             | > In chess, we evaluated PoG against Stockfish 8, level 20
             | [81] and AlphaZero. PoG(800, 1) was run in training for 3M
             | training steps. During evaluation, Stockfish uses various
             | search controls: number of threads, and time per search. We
             | evaluate AlphaZero and PoG up to 60000 simulations. A
             | tournament between all of the agents was played at 200
             | games per pair of agents (100 games as white, 100 games as
             | black). Table 1a shows the relative Elo comparison obtained
             | by this tournament, where a baseline of 0 is chosen for
             | Stockfish(threads=1, time=0.1s).
        
             | ahefner wrote:
             | I'd be interested to see that benchmark. A ~3 GHz Pentium 4
             | sounds like a good reference point for single threaded
             | performance since it's a reasonably modern OoO
             | microarchitecture and reflects the moment that clock
             | scaling stopped.
        
             | Vetch wrote:
             | This sort of evasiveness around speaking on method
             | limitations, down playing or de-emphasizing related work
             | but boosting senior authors previous work is standard
             | academic fare. It's partly a strategy against novelty
             | nitpickers and results in a net negative for all.
             | 
             | I also suspect part of the reason they chose Stockfish 8
             | was as a basis of comparison with AlphaZero. Their
             | baselines for Go and poker are also pretty weak so their
             | emphasis is clearly on displaying generality and reduced
             | domain specialized input, not supremacy.
             | 
             | A single algorithm to play perfect and imperfect
             | information games is difficult to achieve. Standard depth
             | limited solvers and self-play RL result in highly
             | exploitable agents. PoG appears to be very strong at Chess,
             | decently strong at Go and decent at Poker (Facebook AI's
             | ReBeL, the strongest prior work in this area, performed
             | better against slumbot). What's unique about PoG is its
             | ability to also play an imperfect information game
             | (Scotland Yard) that has many rounds and a relatively long
             | horizon (although it still has scaling issues).
        
             | ska wrote:
             | > An analogous situation
             | 
             | It really isn't though. Technical papers have conventions,
             | and they following them reasonably. You expect the methods
             | description to be specific, the abstract not to be
             | hyperbolic, and conclusions to be balanced. The general
             | discussion parts are just that, general.
             | 
             | In the methods area they discuss the exact versions and
             | parameters used, and how they compared them.
             | 
             | In the conclusions:                 In the perfect
             | information games of chess and Go,PoG performs at the level
             | of human experts or professionals, but can be significantly
             | weaker than specialized algorithms for this class of games,
             | like AlphaZero, when given the same resources.
             | 
             | It would have perhaps been interesting to include a more
             | recent stockfish, but it wouldn't really impact the paper.
        
         | dontreact wrote:
         | The name of the game here is generality. For a really general
         | agent, they are looking to have superhuman performance, not get
         | state of the art on every individual task. Beating stockfish 8
         | convinces me that it would be superhuman at chess.
        
           | remram wrote:
           | They could still be honest that it's Stockfish 8, not the
           | Stockfish everyone has. Your product having genuine value
           | does not excuse lying about that value.
        
             | Skyy93 wrote:
             | I observed this kind of behavior in many papers nowadays.
             | This extremely painful for research, because some better
             | candidates could be overseen and FAANG publishs a majority
             | in the ML-paper section. Its a mess.
        
             | ShamelessC wrote:
             | They were? They say they use Stockfish 8 the very first
             | time they mention it.
        
               | hesperiidae wrote:
               | Yup, "In chess, we evaluated PoG against Stockfish 8,
               | level 20 [81] and AlphaZero."
        
               | remram wrote:
               | First time they mention it is page 10:
               | 
               | > one of the strongest and most widely-used programs is
               | Stockfish [81].
               | 
               | Here's the citation, note the date:
               | 
               | > [81] The Stockfish Development Team. Stockfish: Open
               | source chess engine, 2021. https://stockfishchess.org/.
               | 
               | They mention the version number only once, further down,
               | and don't point out that it's out of date since February
               | 2018. All other 11 mentions of it don't have the version
               | number, like in that sentence:
               | 
               | > In Chess, PoG(60000,10) is stronger than Stockfish
               | using 4 threads and one second of search time.
        
               | hesperiidae wrote:
               | >First time they mention it is page 10:
               | 
               | Yeah, so it is! I guess I ran into the same weirdness as
               | ShamelessC, since when I first Ctrl-F:ed the PDF, hit
               | 1/11 was on page 11. Now that I try my damndest to
               | reproduce it, I get 12 hits and the first is that one on
               | page 10.
        
         | david_draco wrote:
         | Isn't the point comparing traditional heuristic techniques
         | against DNN-learned techniques? I understand the latest
         | Stockfish is etching quite close to AlphaZero techniques, but
         | maybe I am wrong.
        
           | tsbinz wrote:
           | It does have the option to use a neural network (nnue) in its
           | evaluation, but it is very different from what AlphaZero/Lc0
           | do. You can choose not to use it, so you still could have a
           | "traditional" evaluation (which would still blow Stockfish 8
           | out of the water). Also, Stockfish 8 isn't the last version
           | before they merged nnue ...
        
         | nixed wrote:
         | the same goes for slumbot in poker, its super old like 2013,
         | the game is played completely different now and current bots
         | would destroy it.
        
           | scrozart wrote:
           | As a commenter above noted, this work is about generality,
           | being able to play every game, and not being the best at
           | every game.
        
             | seoaeu wrote:
             | The abstract claims they beat the "strongest openly
             | available agent in heads-up no-limit Texas hold'em poker".
             | To a non-expert that certainly sounds like they're claiming
             | to be the best
        
               | antonvs wrote:
               | "Openly available" is a strong constraint that's
               | mentioned explicitly.
        
             | Skyy93 wrote:
             | As noted before, the reason for including old tech is to
             | look better. Why not mention the current state of the art
             | and show that with a general player we can come close to
             | this results?
             | 
             | This is just benchmark cherry picking and does not reflect
             | real performance or comparison.
        
           | bluecalm wrote:
           | The problem with poker is that there is money to be made from
           | having a strong AI so there is 0 incentive to release it.
           | What's publicly available are solvers (which solve game
           | abstractions similar to the full game but don't play
           | themselves) and shitty bots.
        
       | fxtentacle wrote:
       | This is a great result, but you can see that it's more of a
       | theoretical case because of this: "converging to perfect play as
       | available computation time and approximation capacity increases."
       | That is true for pretty much all current deep reinforcement
       | learning algorithms.
       | 
       | The practical question is: How much computation do you need to
       | get useful results? Alpha Go Zero is impressive mathematics, but
       | who is willing to spend $1mio daily for months to train it?
       | IMPALA (another Google one) can learn almost all Atari games, but
       | you need a head node with 256 TPU cores and 1000+ evaluation
       | workers to replicate the timings from the paper.
        
         | sillysaurusx wrote:
         | You often don't need anywhere near the amount of compute in
         | these papers to get similar performance.
         | 
         | Suppose you're a business that needs to play games. Most people
         | seem to think that it's a matter of plugging in the settings
         | from the paper, buying the same hardware, then clicking a
         | button and waiting.
         | 
         | It's not. The specific settings matter a lot.
         | 
         | But my main point is that you'll get most of your performance
         | pretty rapidly. The only reason to leave it running for so long
         | is to get that last N%, which is nice for benchmarks but not
         | for business.
         | 
         | DeepMind overspends. Actually, they don't; they're not paying
         | anywhere close to the price of a 256 core TPU. (Many external
         | companies aren't, either, and you can get a good deal by
         | negotiating with the Cloud TPU team.)
         | 
         | But you don't _need_ a 256 core TPU. Lots of times, these
         | algorithms simply do not require the amount of compute that
         | people throw at the problem.
         | 
         | On the other hand, you can also usually get access to that kind
         | of compute. A 256 core TPU isn't beyond reach. I'm pretty sure
         | I could create one right now. It's free, thanks to TFRC, and
         | you yourself can apply (and be approved). I was.
         | https://sites.research.google/trc/
         | 
         | It kills me that it's so hard to replicate these papers, which
         | is most of the motivation for my comment here. Ultimately,
         | you're right: "How much compute?" is a big unknown. But the
         | lower bound is much lower than most people realize (and most
         | researchers).
        
           | fxtentacle wrote:
           | My personal experience was the opposite. I'm currently trying
           | different approaches for building a Bomberman AI for the
           | Bomberland competition that was discussed here on HN a few
           | weeks ago.
           | 
           | "IMPALA with 1 learner takes only around 10 hours to reach
           | the same performance that A3C approaches after 7.5 days."
           | says the paper, but I can run A3C on a cheap CPU-only server
           | but to get that IMPALA timing, I need to spend a lot of
           | money. But my biggest roadblock so far is that I need compute
           | far exceeding what the papers claim.
           | 
           | The diagrams for IMPALA show good performance starting at 1e8
           | environment frames and excellent performance at 1e9 frames.
           | By now, I'm at 2.5e9 frames and performance is still bad. In
           | my opinion, the reason is that the sequence lengths for
           | Bomberland are quite long. To clear a path, you place a bomb,
           | wait 5 ticks for it to become detonatable, then detonate it,
           | then wait 10 ticks for the fire to clear. With 7 possible
           | actions per tick, the chance of randomly executing this 17
           | tick sequence becomes (1/7)^17 = 4e-15. If I calculate
           | optimistically that all moves are valid, too, while we wait,
           | then I can get up to (1/7) _(5 /7)^5_(1/7)*(5/7)^10 = 1e-4.
           | But that still means that at 1e8 env steps, I only have 1000
           | successful executions to learn from.
        
             | ericd wrote:
             | Hm not an expert in this, but would something with a world
             | model help, rather than depending on stochastic random
             | action choices? It seems like it should be possible to
             | learn that a frame sequence where you've been next to a
             | bomb for 6 ticks is rapidly decreasing your expected score,
             | and that your score would be significantly better if you
             | weren't in line with the bomb pretty soon.
        
               | fxtentacle wrote:
               | I'm in the process of attempting just that, with limited
               | success. In my case, I trained a classifier that takes
               | the current surroundings of the player unit and tries to
               | predict that we'll gain an advantage in this segment of
               | the game. I split the game into segments based on when
               | the HP relationships between teams change. And gaining an
               | advantage then means that you take more HP from the enemy
               | team than what you and your teammates lost.
               | 
               | The classifier has on average 90% accuracy which seems
               | good. I then use the likelihood predicted by this
               | classifier to compute the weight with which I want to
               | train each action and if I want to train it positively
               | (by pulling its likelihood of being chosen up) or
               | negatively (pushing the likelihood of that action down).
               | 
               | However, what this model cannot correctly represent is
               | the fact that whether or not a given situation will turn
               | out to be good or bad in the long term is highly
               | dependent on how you play. So if I train this with replay
               | data, I will score the situations in relation to how well
               | those (outdated) AIs could take advantage of them.
               | 
               | Next up, I'll try to fix this issue by introducing a
               | graph-like stochastic structure. The basic idea is that I
               | encode "from this state S if I take action A, then I can
               | reach state T with P percent likelihood" into yet another
               | neural network. If I then identify a state which is
               | really beneficial in the sense that I can reliably
               | convert it into an advantage, then I can use this graph
               | to back-propagate that knowledge so that I get "from this
               | state S, action A takes me to state T, then action B
               | takes me to state U, and U is great".
               | 
               | That should allow me to train with historical data to
               | identify which transitions are possible, and then I can
               | combine that with realtime data about the desirability of
               | each state. So basically I'd do A* pathfinding over the
               | graph of possible states to identify which actions are
               | needed to bring me from my current situation into the
               | closest "I will surely win" situation. Except that the
               | graph is memorized by an AI because the real state-space
               | is huge: 15x15 fields with 6 units + 5 environment states
               | => roughly 11^(15*15) states
        
             | iwd wrote:
             | Not an expert, but I believe many papers on other video
             | games make a single decision for the next X frames at once,
             | possibly including a delay factor that governs exactly when
             | to act. I think OpenAI's Dota2 agent does this.
        
               | fxtentacle wrote:
               | I have experimented with that, too, but in my case it
               | also multiplies the number of potential actions. If I
               | have 7 actions per timestep, grouping them into
               | 3-timestep blocks means I now have 7 _7_ 7 = 343
               | possibilities to choose from.
               | 
               | From what I understand, the OpenAI Dota 2 AI has a long-
               | term strategy module which was mostly trained by
               | imitating 60,000+ replays played by human professional
               | teams. My problem with doing that for the Borderland
               | competition is that I don't have any data source for
               | replays of someone playing the game really well. You
               | control 3 units simultaneously and it's 2 teams against
               | each other, so I'd need 6 dedicated volunteers playing
               | the game for many hours to create a reasonably-sized
               | corpus of human replays. And who says that those people
               | are good at it?
        
             | Javantea_ wrote:
             | I don't have a lot of experience with IMPALA, but the
             | sequence of events you describe should be very easy for an
             | end-to-end system. Assuming you don't have an end to end
             | system, just getting a gradient would result in rapid
             | learning of that sequence. I'm surprised that at 2.5e9
             | frames you're not done. Perhaps there is a hyperparameter
             | issue. Sorry I can't help but it sounds like you are in the
             | same place I am with ML project. Good luck.
        
           | loxias wrote:
           | My thoughts, not being in the field, are parallel to the
           | parent post. "It's nice and all that we're achieving better
           | and better computer performance at things that used to
           | require the human brain, but it seems we're doing so by
           | building larger and larger computers."Not to detract from
           | that achievement, I love large computers in their own right!
           | 
           | I'm a dabbler in Go, and "somewhere below professional" at
           | the game of poker. I've followed the advances in the latter
           | for more than a decade, eagerly reading every paper the CPRG
           | publishes. They use a LOT of compute power!
           | 
           | I know from experience that "The specific settings matter a
           | lot.". For several years, I made my living "implementing
           | papers for hire". It's real work, no argument there.
           | Sometimes the settings _are_ the solution, and heck,
           | sometimes the published algorithm is outright wrong, and you
           | only discover so when trying to implement it.
           | 
           | But the second part of your point, that it's not simply
           | achieving more performance by throwing more transistors at
           | it, I don't have experience with, and I sorta don't believe
           | you. :)
           | 
           | Your comment is quite well written, making me (irrationally?)
           | predisposed to suspect you're correct on factual matters, or
           | at least more of a domain expert than I. Can you cite
           | sources, or simply elaborate more?
        
             | fault1 wrote:
             | > "The specific settings matter a lot.".
             | 
             | Yes, and in the case of deep RL, the ability to to get
             | "lucky" random initialization seems to (still) matter a
             | lot.
             | 
             | I work in real time control systems, which are roughly
             | decision making under uncertainty problems. A lot of the RL
             | research has become noise buoyed with large marketing
             | budgets.
        
         | gwern wrote:
         | > That is true for pretty much all current deep reinforcement
         | learning algorithms.
         | 
         | Is that true? I was unaware that PPO, SAC, DQN, Impala,
         | MuZero/AlphaZero etc would all automatically Just Work(tm) for
         | hidden information games. Straight MCTS-inspired algorithms
         | seem like they'd fail for reasons discussed in the paper, and
         | while PPO/Impala work reasonably well in DoTA2/SC2, it's not
         | obvious they'd converge to perfect play.
        
           | fxtentacle wrote:
           | You can mathematically prove for a lot of different
           | algorithms (including PPO, DQN, IMPALA) that given enough
           | experience with the game world, they will eventually converge
           | to the optimal policy. It's just that the "enough experience"
           | part might be so large that it's practically useless.
           | 
           | If I remember correctly, the DeepMind x UCL RL Lecture Series
           | proves the underlying Bellman equation in this video:
           | https://www.youtube.com/watch?v=zSOMeug_i_M
           | 
           | As for "hidden information" games, I thought the trick was to
           | concatenate the current state with all past states and treat
           | that as the new state, thereby making it an MDP.
        
       | captn3m0 wrote:
       | If you are interested in this, I maintain a list of boardgame-
       | solving related research at
       | https://github.com/captn3m0/boardgame-research, with sections for
       | specific games.
       | 
       | This looks really interesting. It would be a good project to test
       | this against a general card-playing framework to easily test it
       | on a variety of imperfect-information games based on playing
       | cards.
        
         | JoeDaDude wrote:
         | Thank you for posting! Maybe you can include the game of Arimaa
         | [1]. Arimaa was designed to be hard(er) for computers and level
         | the playing field for humans. Algorithms were developed
         | eventually, though I have not kept up to know where that stands
         | today.
         | 
         | [1]. https://en.wikipedia.org/wiki/Arimaa
        
           | captn3m0 wrote:
           | Arima has enough research that it's covered in the Wikipedia
           | section[0] as well as the Chess Programming Wiki[1], which is
           | linked in the README. I'm specifically trying to collect
           | research on contemporary games, which are not so easily
           | available. Chess/Go and alike games are very covered already,
           | however imperfect information games are much rarer for eg.
           | 
           | [0]: https://en.m.wikipedia.org/wiki/Computer_Arimaa
           | 
           | [1]: https://www.chessprogramming.org/Arimaa
        
         | fho wrote:
         | I tried my hand once or twice at (re-)implementing board games
         | [0], so that I could run some common "AI" algorithms on the
         | game trees.
         | 
         | What tripped me up every time is that most board games have a
         | lot of "if this happens, there is this specific rule that
         | applies". Even relatively simple games (like Homeworlds) are
         | pretty hard to nail down perfectly due to all the special
         | cases.
         | 
         | Do you, or somebody else, have any recommendations on how to
         | handle this?
         | 
         | [0] Dominion, Homeworlds and the battle part of Eclipse iirc.
        
           | nicolodavis wrote:
           | You could consider using a library like boardgame.io for
           | this.
        
             | fho wrote:
             | I'll look into that.
        
           | LeifCarrotson wrote:
           | > What tripped me up every time is that most board games have
           | a lot of "if this happens, there is this specific rule that
           | applies". Even relatively simple games (like Homeworlds) are
           | pretty hard to nail down perfectly due to all the special
           | cases.
           | 
           | The key is to build a data-driven state machine, rather than
           | writing logic with a bunch of 'if' statements.
        
             | fho wrote:
             | I am "camp Haskell", so my approach was pretty much data-
             | driven. But what is a state machine if not a big nest of
             | if-else statements? :-)
        
           | captn3m0 wrote:
           | +1 to boardgame.io. It provides very good abstractions for
           | turns, phases, players, and partial information. I've
           | implemented small games with a few hours of effort, and that
           | includes a UI.
        
             | penteract wrote:
             | It's a good set of abstractions, but I've found that the
             | system used for immutability (immerjs) carries noticable
             | performance costs (a factor of more than 2), to the point
             | that it was faster to make a mutable copy of almost all the
             | gamestate at the start of the 'apply move' code.
        
           | iwd wrote:
           | If you're doing it for fun, one option is to start with a
           | simplified version of the game. It's faster to implement and
           | faster to run. And you'll get insights you can apply to the
           | full game.
           | 
           | That's what I did when I applied RL to Dominion, because the
           | complexity of the game depends heavily on the cards you
           | include! See part 3 of https://ianwdavis.com/dominion.html
        
           | anonymoushn wrote:
           | Dominion and Homeworlds are pretty complicated! Maybe you can
           | start with a simpler game like Splendor.
           | 
           | In my 2-player Splendor rules engine, the following actions
           | are possible:
           | 
           | 1. Purchase a holding. (90 possible actions, one for each
           | holding)
           | 
           | 2. If you do not have 3 reserved cards, reserve a card and
           | take a gold chip if possible. (93 possible actions, one for
           | each holding and one for each deck of facedown cards)
           | 
           | 3. If there are 4 chips of the same color in a pile, take 2
           | chips of that color. (5 possible actions)
           | 
           | 4. Take 3 chips of different colors, or 2 chips of different
           | colors if only 2 are available, or 1 chip if only 1 is
           | available. (25 possible actions)
           | 
           | 5. If after any action you have at least 11 chips, return 1
           | chip. (6 possible actions which are never legal at the same
           | time as any other actions)
           | 
           | This still doesn't correctly implement the rules though. In
           | the actual game, you'd be allowed to spend gold chips when
           | you don't need to, which would make purchasing holdings
           | contain extra decisions after you pick which holding to
           | purchase about which chips you'd like to keep.
        
             | fho wrote:
             | I actually played Splendor for the first (three) time(s)
             | some time ago and honestly didn't really like it. It's a
             | very simple game, true. I feel like there are not many
             | decision points for me as a player and therefore there is
             | not much strategy involved. But maybe that is just my view
             | after very few games.
             | 
             | (At the same time that probably makes it a good choice for
             | a game implementation)
             | 
             | Thing is that for all my examples above I had a "good"
             | reason to implement that specific game:
             | 
             | 1. _Dominion_ (shortly after it came out) To evaluate
             | strategies to best my friends (obviously). 2. _Eclipse_ Has
             | a nice rock-paper-scissors type of ship combat, where you
             | can counter every enemy build (if you have enough time and
             | resources). Calculating the odds of winning would be
             | interesting. 3. _Homeworlds_ Seems to be a very fascinating
             | game. But without any players to compete with [0] ... AI to
             | the rescue ;-)
             | 
             | [0] I am aware of SDG where I could play online, but that
             | site is in decay mode. Getting an account involved mailing
             | the maintainer and those times I tried to start a game no
             | players showed up.
        
               | henshao wrote:
               | I think splendor gets more interesting if your opponents
               | are also trying to be strategic. You can see what color
               | chips they are picking up, which lets you know what they
               | are aiming for, which influences what card you want to
               | aim for or reserve. Mid game, you can see what colors
               | other people are missing and try to corner colors to give
               | you room to breath and pick up cards. You can also see
               | the set of colors people are holding to see which of the
               | 4 final bonus point cards are being fought over.
               | 
               | I like the game for what it is. I'd say, surprisingly
               | strategic.
        
               | piyh wrote:
               | It's like if you turned tuning magic mana bases into a
               | stand alone game.
        
         | majani wrote:
         | Imperfect information games will always have a luck element
         | that gives casual players an edge. That's basically the appeal
         | of card games over board games.
        
           | ketzo wrote:
           | And why so many board games incorporate decks/hands of cards.
        
           | mathgladiator wrote:
           | Not just luck but deception as well which takes some games to
           | new levels.
        
         | alper111 wrote:
         | This looks very good, thanks.
        
       | antonpuz wrote:
       | Anyone knows whether the agent is publicly available?
        
       | [deleted]
        
       ___________________________________________________________________
       (page generated 2021-12-08 23:03 UTC)