Writing the test scene (the opening scene) for All Systems Blue -- the space opera. Which means discovering the character's voice, the style in which I am writing, how well free indirect speech in third person is working for me and how I mean to handle it, how the world is presented, and as expected; I'm running into a lot of world-building questions.
You really don't know what you need until you've got people in it, regardless of whether you are a games master with a bunch of unruly players, or the writer and sole creative spirit.
I'm aiming for the 90-100K range with this one.
Also dealing with cover artists and editors and blurb-wranglers and I'm sick of it. I just want to put The Early Fox out (the archaeological-adventure series which is currently getting staged as a mystery-thriller) and return my concentration to writing.
I feel like half these people at Fiverr and in the cover shops are not actually talking to me, but shoving all that correspondence through Claude. I'll say in my query that I need help with cover art on a novel, and I'll get back an ELIZA-like "I understand you want to talk about cover art." It really is that blatant.
My current -- subject to change at any moment -- feeling about AI is that it is a tool and it is inevitable. We are dealing with the teething problems of any new technology now, and much of the presentation is oversold and much of the adoption is premature. And that means the problems are also significant right now.
I won't say there aren't systemic issues but there are with all technologies. There doesn't seem to be anything particularly new here. We humans have always been prone to offloading whatever we can. Saving labor. Saving time by having something else do it, or by cutting corners when we can get away with it. And offloading the need to think about everything in our environment, relying on simple rules or better yet, systems of rules (like religions or political beliefs) that give answers we can use and more-or-less rely on.
Right now, the market pressures are extreme and the pace of economic activities (and change) is extremely rapid. That's an opening for scammers, charlatans, and for human laziness to overtake any spare brain cells we might have to consider whether the AI is giving us good information.
Eventually it will be subtler. The tendencies of AI will be as well understood as how watercolor paints work or what you can realistically accomplish in three weeks of rehearsal. Implicit there is that the results will often be good enough, but there will exist a boundary where it isn't good enough and it will be considered the wrong tool. Going along with this, AI becomes more interactive. I am seeing this in AI video generation, where it has moved sharply beyond where the user puts in a prompt and hopes, to where the user can manage at a closer scale, dictating camera moves, pre-designing characters, and so on.
The earlier version of it was a black box, and this is how it is still being used by the terminally lazy. Used intelligently and with a good work ethic, yes, it can help edit a novel. Used by someone on Fiverr who dares to call themselves an editor and might not have even looked at the manuscript themselves? That's crap, that's insulting, and I'm damned tired of people who think I'm sucker enough to keep paying them for it.
Because even when you toss out the stealing and the environmental costs and the many and fascinating abuses, even when you boil it down to just whether the result is good... the result ain't. It can be, but when it is used as a crutch, when it is used out of laziness, when it is used only for quick cash, those results are not.
Anyhow.
Not one to stow thrones, I've been messing around with LTX far too much lately. I haven't installed some of the front-end stuff that makes it more controllable, as the Comfy-UI front end has a bad habit of breaking everything with each mandatory install (its the freeware version of Windows 11!) To protect my rig I leave it off line. And that makes it harder to install new tricks.
The latest down the pipe is MinMax H3, which is even more scriptable with multiple references of multiple types. And much higher character consistency. LTX 2.3 has that same object permanence problem all AI has: you can start with a fully rendered character, but if they ever leave the frame, what will return may be something different. So does MinMax, but at least you can give it a stable reference to look back on.
Having a combined video-audio output is a mixed blessing. Mostly to the good, because character animation matches the vocal track in a way even the audio-track-plus-starting-image method in Wan2.2 didn't. The weird thing is, since you are giving it a starter image (the usual approach) you can make multiple clips with character consistency, but you won't get consistent vocal performance. That makes it very difficult to marry clips together into a longer video.
It also doesn't normally loop, though there are workflows to achieve that. My last run of experiments was mostly pressure-testing by running either pixel size or frame length out to excess numbers to see what would happen. I could get one or two renders out to 1400 frames -- almost a minute at the default frame rate -- before memory leaks started throwing errors. And at a much much higher consistency than anything else I'd tried before.
The other thing I was messing around with was multi-character dialogue. It did sometimes lose track of who was who, and the range of performance styles can be limited, but it was pretty surprising how natural the results could get. In fact, I just finished an experimental render in which I prompted a 1970s holdout to stroll into the scene carrying a guitar and sing a song badly. First try, and it did pretty much everything I'd been hoping.
Which is another thing. Although I'm told MinMax H3 is even better at smoothly combining character action, dialogue, and camera movements, I've been impressed at how much more agile LTX 2.3 is about camera moves. It was pulling teeth to get Wan 2.2 to even pull a zoom. LTX 2.3 seems to understand the dramatic intent of a camera move, and will integrate it with action.
Of course, it probably helps that I'm writing camera motions and character actions in between chunks of dialogue. As with all things AI, it doesn't always do them in the order prompted. But quite often it comes close to story-telling norms.
Now if only I could do multiple shots and cut them together. I still haven't done a character LoRA and there's an additional step of converting a LoRA to work with LTX 2.3 (and I really miss the LoRA I had with WAN 2.2). So maybe I will try that MinMax install.
It seem silly making video clips nobody will ever see. But some days I just want to sit down in front of the gaming rig and mess around... and how different is this from building factories nobody will ever visit?
No comments:
Post a Comment